Technology
Cognition launches SWE-2 coding model built on Kimi K33 base
Cognition has released its SWE-2 coding model to Devin users, according to Hacker News.
Latest update: Benchmark disclosures reveal SWE-2 scored 27.3% on Terminal-Bench 4.0, trailing rival models Fable 5.1 and GPT-6 Astra.
The short version
- Cognition introduced SWE-2, a model post-trained on the 2.8-trillion-parameter Kimi K33 base using reinforcement learning.[Hacker News]
- The developer reports the model scored 50.0% on FrontierCode 1.1 Main at 64% lower cost than Fable 5.1, but reached only 27.3% on Terminal-Bench 4.0.[Hacker News · Hacker News]
- SWE-2 is available in Devin Desktop and CLI, with Web and Fusion rollouts underway, but weights and public per-token APIs remain unreleased.[Hacker News · Hacker News]
- Reported efficiency and performance claims stem entirely from vendor benchmarks and await independent replication.[Hacker News]
Key facts
- Cognition post-trained SWE-2 on the 2.8-trillion-parameter Kimi K33 base model, scaling reinforcement learning into the multi-trillion-parameter regime.[Hacker News]
- Cognition claims SWE-2 reached 50.0% on FrontierCode 1.1 Main while costing 64% less than Fable 5.1.[Hacker News · Hacker News]
- SWE-2 scored 27.3% on Terminal-Bench 4.0, trailing significantly behind Fable 5.1's 55.8% and GPT-6 Astra's 57.9%.[Hacker News · Hacker News]
- SWE-2 is accessible in Devin Desktop and CLI with deployments underway for Web and Fusion, but lacks public model weights or a standard per-token API.[Hacker News · Hacker News]
What remains uncertain
- All published benchmark results and relative cost efficiencies are self-reported by Cognition and remain unverified by third-party evaluators.[Hacker News]
Sources
Outlet counts describe coverage, not independent confirmation. Reports may share a wire service or original source.