Technology
Artificial Analysis updates Intelligence Index with new benchmarks and private data
Anthropic and OpenAI top the revised rankings as evaluations shift toward realistic tasks and anti-gaming controls.
The short version
- Artificial Analysis launched Intelligence Index v4.2, adding the AA-Briefcase and GDP.pdf benchmarks while removing the saturated GPQA Diamond.[Hacker News]
- Anthropic's Claude Fable 5.1 leads the updated overall index, followed closely by OpenAI's GPT-6 Astra, with Meta ranked third.[Hacker News]
- Private, held-out test sets now represent 40% of the index weighting to prevent labs from gaming evaluation metrics.[Hacker News]
- The organization plans further incremental updates ahead of a full v5 release.[Hacker News]
Key facts
- Artificial Analysis announced Intelligence Index v4.2 on September 4, 2026, as an interim update ahead of an upcoming v5 release.[Hacker News]
- The update incorporates the AA-Briefcase and GDP.pdf evaluations while dropping the saturated GPQA Diamond benchmark.[Hacker News]
- Private, held-out test sets were increased to account for 40% of the total index weighting to limit gaming by model developers.[Hacker News]
- Anthropic's Claude Fable 5.1 took the top spot on the overall index, followed by OpenAI's GPT-6 Astra in second and Meta in third.[Hacker News]
- OpenAI's GPT-6 Astra led the GDP.pdf document reasoning benchmark with a 33.2% score.[Hacker News]
Sources
Outlet counts describe coverage, not independent confirmation. Reports may share a wire service or original source.
- Artificial Analysis Intelligence Index v4.2Hacker News