Technology
EEBench releases benchmark evaluating AI models on circuit board design
A new code- and simulation-based evaluation tests how frontier artificial intelligence models navigate circuit design, component selection, and tolerance constraints.
The short version
- EEBench has introduced a deterministic benchmark measuring AI performance on analog and digital circuit design using declarative code and SPICE simulations.
- Anthropic's Claude Opus 5 led tested models with a 61.6% score, followed by Grok 4.6 and Claude Fable 5.1, while tested OpenAI models scored under 45%.
- The evaluation tests requirements, design, and simulation verification, but does not yet assess physical board layout, fabrication, or product bring-up.
Key facts
- EEBench V1 tests AI models on 13 analog and digital circuit design tasks using the atopile declarative code language and deterministic SPICE simulations.[Hacker News]
- Claude Opus 5 achieved the top score of 61.6%, while Claude Fable 5.1 scored 56.4%.[Hacker News]
- Grok 4.6 scored 57.1% on standard testing and reached 60.0% with high reasoning effort in data reported in xAI's model card.[Hacker News]
- OpenAI's tested models scored lower, with GPT-5.5 reaching 42.3% and GPT-5.6 Sol reaching 39.4%.[Hacker News]
- The benchmark grades technical performance across component tolerances and real datasheet specifications, factoring in bill-of-materials cost efficiency once designs function.[Hacker News]
What remains uncertain
- How newer models like OpenAI's GPT-6 Astra or the unreleased Grok 4.7 will perform on the benchmark remains unknown because they have not yet been evaluated.[Hacker News]
- Whether AI models can handle physical board layout, manufacturing pipelines, and full hardware bring-up remains unmeasured by the current version of the benchmark.[Hacker News]
Sources
Outlet counts describe coverage, not independent confirmation. Reports may share a wire service or original source.
- Can AI design circuit boards yet?Hacker News