September 6, 2026BenchmarkAgents

Can AI Design a Circuit Board? EEBench Says 61.6% of One

The team behind atopile built and funded EEBench, a benchmark asking whether AI can actually design electronics, and published first results that hit the HN front page at 352 points (https://eebench.org/blog/can-ai-design-circuit-boards-yet/). The answer: Claude Opus 5 scores 61.6%, Grok 4.6 gets 57.1%, Claude Fable 5.1 sits at 56.4%, and OpenAI trails badly — GPT-5.5 at 42.3%, GPT-5.6 Sol at 39.4%.

The design choice that makes this benchmark work: no GUI CAD. Everything runs through atopile's declarative code, so agents manipulate components, connections and electrical constraints as text, then get judged by simulation. The 13 tasks in V1 are unglamorous and real — a residential energy meter hold-up circuit where you pick capacitors under real-world tolerances, multiple-feedback low-pass filters where component ratios must be synthesized. Grading covers manufacturer datasheets, tolerance corners, cost efficiency and parts availability.

The headline finding cuts both ways. Models know far more electronics than their usual tool-mediated output suggests — free them from clicking through CAD and the knowledge shows up. But physics keeps catching them: designs failed on insufficient effective capacitance even when the nominal math looked right, the exact class of mistake that separates a textbook from a shipped board.

Two things worth noticing. Opus 5 beating the newer Fable 5.1 suggests deep domain knowledge lives in the big slow models, not the agentically-polished ones. And benchmarks are marching from software into atoms — after terminals, browsers and science, it is now tolerance corners and supply chains. The gap between 61.6% and a board you would actually manufacture is where the next two years of this thread live.
← Previous
Spotify Cut Claude Code Tokens 90% by Giving It a Cheap Intern
Next →
IBM Bob Walks Into the Coding-Agent Party Carrying a Mainframe
← Back to all articles

Comments

Loading...
>_