← Latest briefing

Technology

Real-SWE benchmark tests frontier artificial intelligence models on private enterprise codebases

Hacker News reports that the evaluation found leading systems struggle to solve most real-world software engineering tasks.

The short version

  • A newly launched benchmark, Real-SWE, assesses artificial intelligence models using licensed enterprise software repositories.[Hacker News]
  • Fable 5.1 led evaluated models with a 38.8% task resolution rate, followed by GPT-6 Astra at 33.8%.[Hacker News]
  • Estimated rollout costs ranged between $2.50 and $6.96, while over 70% of runs failed regardless of duration.[Hacker News]

Key facts

  • Real-SWE is an evaluation framework testing frontier models on private, enterprise codebases licensed from companies.[Hacker News]
  • Fable 5.1 secured the top resolution score at 38.8%, followed by GPT-6 Astra at 33.8% and Gemini 3.8 Flash at 31.2%.[Hacker News]
  • Across evaluated systems, estimated rollout expenses ranged from $2.50 to $6.96.[Hacker News]
  • Failures remained consistent across durations, with 71.4% of rollouts under 10 minutes failing compared to 73.4% of longer attempts.[Hacker News]

Sources

Outlet counts describe coverage, not independent confirmation. Reports may share a wire service or original source.