← Latest briefing

Technology

Stanford and partners introduce Terminal-Bench-Science to test AI agents on scientific research

The continuous benchmark measures AI model performance across 70 scientific workflows, with top models resolving under one-third of tasks.

The short version

  • Researchers at Stanford University and partner institutions released Terminal-Bench-Science 0.1 to evaluate AI agents on real-world scientific research workflows.
  • The initial release includes 70 curated tasks across five scientific domains, selected from 920 initial proposals.
  • The top-performing model, Claude Opus 5, achieved a 30% resolution rate, while the strongest open model reached 8.1%.
  • Organizers are currently developing Terminal-Bench-Science 0.2, setting an October 5, 2026, deadline for new task contributions.

Key facts

  • Terminal-Bench-Science 0.1 was developed by researchers at Stanford University alongside the Terminal-Bench team and domain experts across multiple scientific disciplines.[Hacker News]
  • The benchmark's initial version features 70 curated tasks spanning the life, physical, Earth, mathematical, and engineering sciences out of 920 submitted proposals.[Hacker News]
  • Claude Opus 5 recorded the highest resolution rate among evaluated models at 30.0%, incurring an evaluation cost of $7,000.[Hacker News]
  • GPT-5.6 Sol and Claude Fable 5 followed in accuracy, achieving resolution rates of 22.4% and 21.4% respectively.[Hacker News]
  • GLM 5.3 placed as the highest-performing open model, registering an 8.1% resolution rate.[Hacker News]
  • The project deadline for submitting pull requests for version 0.2 of the benchmark is set for October 5, 2026.[Hacker News]

What remains uncertain

  • Future performance trends across emerging models remain subject to evaluation as new tasks and updated frontier models are introduced in subsequent releases.[Hacker News]

Sources