Technology
Real-SWE benchmark tests frontier artificial intelligence models on private enterprise codebases
Hacker News reports that the evaluation found leading systems struggle to solve most real-world software engineering tasks.
The short version
- A newly launched benchmark, Real-SWE, assesses artificial intelligence models using licensed enterprise software repositories.[Hacker News]
- Fable 5.1 led evaluated models with a 38.8% task resolution rate, followed by GPT-6 Astra at 33.8%.[Hacker News]
- Estimated rollout costs ranged between $2.50 and $6.96, while over 70% of runs failed regardless of duration.[Hacker News]
Key facts
- Real-SWE is an evaluation framework testing frontier models on private, enterprise codebases licensed from companies.[Hacker News]
- Fable 5.1 secured the top resolution score at 38.8%, followed by GPT-6 Astra at 33.8% and Gemini 3.8 Flash at 31.2%.[Hacker News]
- Across evaluated systems, estimated rollout expenses ranged from $2.50 to $6.96.[Hacker News]
- Failures remained consistent across durations, with 71.4% of rollouts under 10 minutes failing compared to 73.4% of longer attempts.[Hacker News]
Sources
Outlet counts describe coverage, not independent confirmation. Reports may share a wire service or original source.