October 7, 2026Open SourceInfrastructureResearch

openTPU: Agents Designed an AI Accelerator, and It Runs Gemma 4 on a $300 Board

A repo called openTPU hit 180 points on Hacker News on Tuesday with a one-line claim: an open-source AI accelerator, developed by AI. The author, who previously used agents to build RISC-V CPU cores in a project called auto-arch-tournament, pointed the same method at inference hardware. The question the README asks is whether agents can design the chip that runs their own inference.

What exists is a complete stack in one monorepo: SystemVerilog hardware design, an instruction set, a bit-exact simulator, a kernel language with its compiler, and host software driving a real PCIe card. The card is a decommissioned data-center FPGA board, an Inspur unit with a Xilinx Kintex-7 and two DDR3 channels that hobbyists buy for around $300. It runs ten modern models with real weights. LFM2.5-230M decodes at 85.8 tokens per second in 4-bit, Qwen3-0.6B at 31.3, Gemma 4 E2B at 12.1, Phi-4-mini at 6.6. Mixture-of-experts models bigger than the card's 4GB run with experts streamed from host storage: Qwen3.5-35B-A3B at 3.95 tokens per second. Every configuration matches the simulator token for token.

The improvement loop is the story. In the HN thread the author says the design started at a few tokens per second and reached 80+ on the smallest models through a recursive self-improvement loop. The README shows what that loop optimizes: DRAM bandwidth utilization went from 82-87% to 91-94% between production images, the memory controller was swapped for LiteDRAM with on-card calibration in 12 seconds, and a tournament of Vivado runs keeps working on timing margin, which currently closes at 133 MHz by 0.032 nanoseconds. Decode is memory-bound, so that bandwidth number is the whole game.

The hardware is tiny and the method is not. A tournament loop with a measurable metric, a simulator that can say pass or fail, and an agent that proposes changes is exactly the shape that worked for kernels and compilers this year. HN's reaction split between "vibe connoisseuring" and the obvious joke about recursive self-improvement being both the warning and the product. The author's position is that chip design tooling has moved more in months than in years. The evidence here is a working card, which is more than most of those claims come with.

Link: github.com/FeSens/openTPU
← Previous
REA Gives Your Agent Ghidra, and GitHub Gave It 3,000 Stars in a Day
Next β†’
ErdΕ‘s Problems Freezes Comments: The AI Proof Flood Pushed Out the Humans
← Back to all articles

Comments

Loading...
>_