GPT-6 Astra Now on Par With Claude Fable 5.1 in Updated Artificial Analysis Index
Set Trending Topics as a preferred source on Google.
Artificial Analysis has just raised its Intelligence Index to version 4.3. Two models now share the top of the updated ranking: GPT-6 Astra (max) from OpenAI and Claude Fable 5.1 (max with fallback) from Anthropic both score 53 points. They are followed by Claude Opus 5 (51), Claude Fable 5 (50), Meta’s Muse Spark 1.3 (48) and GPT-5.6 Sol (47). Among open weights models, GLM-5.3 from Z.ai and Kimi K3 from Moonshot lead the field, ahead of GLM-5.3-Flash (42), Qwen3.8 2.4T A95B (40) and DeepSeek V4 Pro (36).
For OpenAI, that amounts to a rapid climb: within a few days and across two updates of the index, GPT-6 Astra has moved from fifth place to first. Less than a week separates the model’s debut from its current position, and in that window Artificial Analysis changed the composition of its test suite twice.
What Changed in the Index
The index still consists of ten individual evaluations, two of which have been replaced or upgraded. Terminal-Bench jumps from version 2.1 to version 4.0 and uses 66 tasks to test whether an agent can complete complex work through the terminal, spanning software, machine learning, science, operations, security, hardware and media. GPT-6 Astra pulls clearly ahead here with 59.1 percent, while Claude Fable 5.1 reaches 52.0 percent, Claude Opus 5 49.0 percent and GPT-5.6 Sol 39.9 percent.
Also new to the index is AutomationBench-AA, which replaces the previous banking evaluation 𝜏³. The benchmark comes from Zapier and covers business processes across finance, HR, marketing, operations, sales and support: 657 tasks in which agents work inside simulated business applications and have to discover the relevant APIs themselves. Any guardrail violation drops that task’s score to zero. GPT-6 Astra leads here too with 68.5 percent, ahead of Grok 4.6 (66.7 percent) and GLM-5.3 (62.2 percent). Astra completes 41.6 percent of workflows in full and without a violation, compared with 32.1 percent for Claude Fable 5.1.
The share of evaluations whose questions or answers stay private rises from 40 to 45 percent. Artificial Analysis wants to make it harder for providers to optimize their models against known benchmarks. Category weights remain unchanged at Agents 30 percent, Coding 20 percent, General 30 percent and Scientific Reasoning 20 percent.
Same Score, Half the Price
The two leaders sit far apart on cost. An average index task runs to 3.26 US dollars with GPT-6 Astra and 7.63 US dollars with Claude Fable 5.1, a gap of 57 percent. The spreads further down the field are similar: GLM-5.3-Flash and GPT-5.6 Terra both score 42 points, yet the Chinese model costs 0.25 US dollars per task against 1.40 US dollars for the OpenAI model. Four labs occupy the Pareto frontier between intelligence and cost, with OpenAI holding most of the curve.
Three Index Versions in One Week
The speed of the shift is as striking as the result itself. When GPT-6 Astra first ran through the index a few days ago, it scored 61 points, exactly level with its own predecessor GPT-5.6 Sol and five points behind Claude Fable 5.1 (66). That put it in the lower half of the leading group, behind several Anthropic models and Meta. Last Friday brought version 4.2 of the index, where Fable 5.1 led with 57 points against Astra’s 55, narrowing the gap to two points. With version 4.3 it has closed entirely.
The models themselves stayed the same throughout. What changed were the measuring instruments. Because the index is recalculated in full after every revision, one model’s rise necessarily pushes others down: between version 4.2 and 4.3, Claude Fable 5.1 loses four points, Claude Opus 5 three, and Muse Spark 1.3 five. That makes it hard for observers to tell whether a change in rank reflects actual capability or the composition of the test set.
Then there is the context in which both updates emerged. After GPT-6 Astra’s debut, the index drew criticism because other evaluations placed the model considerably higher: Epoch AI ranked Astra first out of 267 models across more than 50 benchmarks, and OpenAI’s own numbers showed a clear lead as well. Artificial Analysis explains the two interim versions as pulling forward elements of the planned index v5 in order to keep pace with the rate of model releases. Each individual change, the company argues, stands on its own merits: closer to real-world tasks, more private test sets, less saturation.
A look at the other major leaderboards shows how differently the model is currently assessed. At Epoch AI GPT-6 Astra sits in first place, based on more than 50 benchmarks. At Arena.ai, where rankings come from blind comparisons made by users, the model carries no rating at all in the text and overall categories, which are led by Claude Fable 5.1 with 1,231 points. So far Astra appears in the arena only in the specialist WebDev category, which it tops with 1,797 points.
Both readings find support in the data. The methodological case rests on the new evaluations being substantially harder, with scores falling across the entire field, which argues against a tweak made to favor a single model. The counterargument is that a ranking rebuilt three times in one week holds little value as a time series.

