Technology
CodeRabbit evaluation finds GPT-6 Astra improves multi-file code review performance
The model achieved higher bug detection gains on complex reviews despite carrying premium API pricing.
The short version
- In a benchmark evaluation by CodeRabbit, OpenAI's GPT-6 Astra caught roughly 4% more labeled bugs than GPT-5.6 Sol and 22% more than Opus 5 overall.[Hacker News]
- Performance improvements were larger in complex cross-file reviews, reaching a 20% advantage over Sol and 33% over Opus 5.[Hacker News]
- Astra carries a premium pricing tier of $10 per million input tokens and $50 per million output tokens, matching Claude Fable 5.1 base rates.[Hacker News]
- CodeRabbit noted that these early benchmark metrics do not guarantee identical defect reduction or performance across every repository pull request.[Hacker News]
Key facts
- CodeRabbit's testing showed GPT-6 Astra caught approximately 4% more labeled bugs through actionable findings than GPT-5.6 Sol and 22% more than Opus 5.[Hacker News]
- In more difficult cross-file evaluations, Astra's advantage rose to 20% over Sol and 33% over Opus 5.[Hacker News]
- Published API pricing for GPT-6 Astra is set at $10 per million input tokens and $50 per million output tokens.[Hacker News]
- Neither CodeRabbit nor its model providers train artificial intelligence models on customer proprietary code or personal data from private reviews.[Hacker News]
- OpenAI supports zero data retention for eligible API customers on Astra, whereas Anthropic maintains a default 30-day retention period on Fable with conditional exceptions.[Hacker News]
What remains uncertain
- CodeRabbit emphasized that the evaluation results represent only a segment of review capability and do not establish overall quality rankings or predict defect rates across all pull requests.[Hacker News]
Sources
Outlet counts describe coverage, not independent confirmation. Reports may share a wire service or original source.