The barrier between software architecture and silicon design is crumbling. FeSens has released openTPU, an open-source AI accelerator explicitly developed by AI agents. This project isn't just a theoretical exercise; it delivers a working SystemVerilog design that runs ten modern language models on physical hardware. The core question driving the repo is pragmatic: can AI agents build the specific chip required to run their own inference workloads efficiently?
Hardware and Performance Benchmarks
The implementation targets an Inspur YPCB-00338 card featuring a Xilinx Kintex-7 xc7k480t FPGA and two DDR3 channels. The results are tangible and measured. For the LFM2.5-230M model using int8 weights, the card achieves 59.0 tokens per second on the device, hitting 85% of the DRAM peak bandwidth. Switching to 4-bit weights with an int8 head boosts this to 85.8 tokens per second. Larger models like SmolLM3-3B and Phi-4-mini (3.8B) run at 5.00 and 3.99 tokens per second respectively in int8, maintaining high DRAM utilization rates near 94%.
A Monorepo for Full-Stack Understanding
openTPU distinguishes itself as a learning resource by consolidating the entire stack into one readable monorepo. Developers can trace the path from a Python matrix multiplication down to the physical wires. The repository includes the hardware design in SystemVerilog, a bit-exact simulator, a custom kernel language with its compiler, and the host software. This transparency allows builders to understand exactly how data moves through the system, avoiding the black-box nature of many commercial accelerators.
Host Integration and Offloading Strategies
The system minimizes host overhead by having the FPGA run a compiled decode program that reads positions from registers and manages its own embedding lookups. For models larger than the card's 4 GiB memory capacity, such as the Qwen3.5-35B-A3B, the design employs a mixture-of-experts streaming strategy. Experts are stored on host storage and copied into the card's DRAM slots as needed. The Qwen3.5-35B model achieves 3.95 tokens per second with 62% of expert uses hitting the local slots, demonstrating that AI-designed hardware can handle complex, large-scale inference tasks.
Key Takeaways
- AI agents successfully designed a functional SystemVerilog accelerator capable of running real-world LLMs.
- The project achieves bit-exact accuracy between the simulator and physical FPGA hardware.
- Full-stack visibility allows developers to debug performance from the ISA down to the RTL.
- Mixture-of-experts models larger than on-chip memory are supported via efficient host-to-card streaming.
The Bottom Line
openTPU proves that AI agents are ready to design their own silicon, but the true value lies in its transparency. By exposing the full stack from kernel to RTL, it turns a niche hardware experiment into a powerful educational tool for developers who want to understand the physical cost of inference.