GPT-6 Still Behind Fable 5.1 As Artificial Analysis Overhauls Intelligence Index
Set Trending Topics as a preferred source on Google.
The most closely watched yardstick for AI models has just been reworked. Artificial Analysis recently released version 4.2 of its Intelligence Index, the aggregate score that labs, investors and developers reach for when they want to place a language model in the wider field. The update pushes the benchmark toward harder, more realistic tasks and toward a much larger share of secret test data. At the top of the leaderboard, though, the order stays as it was: Anthropic’s Claude Fable 5.1 leads, and OpenAI’s GPT-6 Astra follows in second place.

An Interim Step Ahead of Version 5
Artificial Analysis describes v4.2 as an interim release. Eight months have passed since Index v4 launched, and the team deliberately held back changes during that period to keep the score stable across several major model releases. Because the frontier has moved so quickly in recent weeks, the company is now pulling forward individual components of the planned version 5. More incremental releases are expected to follow.
What Is New in the Index
Two benchmarks join the Index, and one drops out.
- AA-Briefcase: The company’s in-house evaluation, built around a private held-out test set, measures models on realistic agentic knowledge work. Industry experts designed multi-week projects made up of many linked subtasks and thousands of source files. Grading combines rubrics with pairwise comparisons along three axes: verifiable task success, analytical quality and presentation quality.
- GDP.pdf: Developed by Surge AI, this test measures single-turn professional document reasoning across 100 PDFs from ten domains. Models have to synthesize evidence spread over 4,592 pages, covering text, tables, charts, footnotes and exclusions. Answers are scored against 1,275 expert-authored atomic criteria, and the headline All-pass Rate credits a task only when every single criterion is met.
- GPQA Diamond: The scientific reasoning benchmark leaves the Index because it has become saturated and now offers little separation between frontier models.
A structural change comes on top of that. Private, held-out test sets now account for 40 percent of the Index weighting, double the figure in v4.1. Alongside AA-Briefcase, that share includes AA-Omniscience and the solutions for CritPt. The goal is to make it harder for labs to optimize directly against the benchmark, and Artificial Analysis says the held-out portion will grow further in version 5.
The grading stack was upgraded as well. AA-LCR now runs at version 1.1 with its own grading system prompt and corrected answer keys, while GDPval-AA v2 and AA-Briefcase received improved sampling and a re-anchored Elo scale so that ratings stay stable as new models are added. For SciCode, the grading sandboxes were made more robust, so that slow but correct code counts as a pass.
Anthropic Ahead, OpenAI Behind
The upper end of the table keeps its familiar shape. Claude Fable 5.1 tops the Index, the model that arrived recently as Anthropic’s most capable and also its pricier offering. GPT-6 Astra sits behind it, gaining four points on its predecessor GPT-5.6 Sol. The gap to the top therefore remains, even though OpenAI positioned the model at launch as a step toward AGI. Astra already landed behind the leading models from Anthropic and Meta in benchmarks right after its release, and the revised methodology leaves that picture intact.
Meta ranks as the third lab on the leaderboard, followed by SpaceXAI, Moonshot with Kimi, Z.AI and Google.
Where Each Model Shines
The individual tests give a more granular view.
- AA-Briefcase: Claude Fable 5.1 and Opus 5 lead here, ahead of GPT-6 Astra and Muse Spark 1.3. Astra improves on GPT-5.6 Sol by roughly 85 Elo points, a substantial jump in agentic knowledge work.
- GDP.pdf: Document reasoning reverses the order. OpenAI leads with GPT-6 Astra at 33.2 percent, GPT-5.6 Sol reaches 28.2 percent, and Claude Fable 5.1 comes in at 26.2 percent. The low absolute numbers show how demanding the all-pass criterion is.
- Cost: Four labs share the updated Cost per Task Pareto frontier: Anthropic, OpenAI, Meta and Z.AI.
- Token efficiency: GPT-6 Astra dominates the output token frontier and works more sparingly than almost every other model near the intelligence frontier. Claude Fable 5.1, Grok 4.5 and Gemini 3.5 Flash-Lite sit at either end of the curve, with models below 25 Index points excluded.
For practitioners, the split is fairly clear. Teams that need agentic project work across many linked steps will find the strongest results at Anthropic today, while teams processing large document sets or watching token costs have a strong case for GPT-6 Astra. Artificial Analysis has published the details of its approach in its benchmarking methodology documentation.

