← Latest briefing

Technology

CodeRabbit evaluation finds GPT-6 Astra improves multi-file code review performance

The model achieved higher bug detection gains on complex reviews despite carrying premium API pricing.

The short version

  • In a benchmark evaluation by CodeRabbit, OpenAI's GPT-6 Astra caught roughly 4% more labeled bugs than GPT-5.6 Sol and 22% more than Opus 5 overall.[Hacker News]
  • Performance improvements were larger in complex cross-file reviews, reaching a 20% advantage over Sol and 33% over Opus 5.[Hacker News]
  • Astra carries a premium pricing tier of $10 per million input tokens and $50 per million output tokens, matching Claude Fable 5.1 base rates.[Hacker News]
  • CodeRabbit noted that these early benchmark metrics do not guarantee identical defect reduction or performance across every repository pull request.[Hacker News]

Key facts

  • CodeRabbit's testing showed GPT-6 Astra caught approximately 4% more labeled bugs through actionable findings than GPT-5.6 Sol and 22% more than Opus 5.[Hacker News]
  • In more difficult cross-file evaluations, Astra's advantage rose to 20% over Sol and 33% over Opus 5.[Hacker News]
  • Published API pricing for GPT-6 Astra is set at $10 per million input tokens and $50 per million output tokens.[Hacker News]
  • Neither CodeRabbit nor its model providers train artificial intelligence models on customer proprietary code or personal data from private reviews.[Hacker News]
  • OpenAI supports zero data retention for eligible API customers on Astra, whereas Anthropic maintains a default 30-day retention period on Fable with conditional exceptions.[Hacker News]

What remains uncertain

  • CodeRabbit emphasized that the evaluation results represent only a segment of review capability and do not establish overall quality rankings or predict defect rates across all pull requests.[Hacker News]

Sources

Outlet counts describe coverage, not independent confirmation. Reports may share a wire service or original source.