Technology
Stanford and partners introduce Terminal-Bench-Science to test AI agents on scientific research
The continuous benchmark measures AI model performance across 70 scientific workflows, with top models resolving under one-third of tasks.
The short version
- Researchers at Stanford University and partner institutions released Terminal-Bench-Science 0.1 to evaluate AI agents on real-world scientific research workflows.
- The initial release includes 70 curated tasks across five scientific domains, selected from 920 initial proposals.
- The top-performing model, Claude Opus 5, achieved a 30% resolution rate, while the strongest open model reached 8.1%.
- Organizers are currently developing Terminal-Bench-Science 0.2, setting an October 5, 2026, deadline for new task contributions.
Key facts
- Terminal-Bench-Science 0.1 was developed by researchers at Stanford University alongside the Terminal-Bench team and domain experts across multiple scientific disciplines.[Hacker News]
- The benchmark's initial version features 70 curated tasks spanning the life, physical, Earth, mathematical, and engineering sciences out of 920 submitted proposals.[Hacker News]
- Claude Opus 5 recorded the highest resolution rate among evaluated models at 30.0%, incurring an evaluation cost of $7,000.[Hacker News]
- GPT-5.6 Sol and Claude Fable 5 followed in accuracy, achieving resolution rates of 22.4% and 21.4% respectively.[Hacker News]
- GLM 5.3 placed as the highest-performing open model, registering an 8.1% resolution rate.[Hacker News]
- The project deadline for submitting pull requests for version 0.2 of the benchmark is set for October 5, 2026.[Hacker News]
What remains uncertain
- Future performance trends across emerging models remain subject to evaluation as new tasks and updated frontier models are introduced in subsequent releases.[Hacker News]