Artificial Analysis

GPT-6 Astra Trails Top Models From Anthropic in Benchmarks

Altman vs. Amodei. © WEF / Collage Trending Topics
Altman vs. Amodei. © WEF / Collage Trending Topics

Set Trending Topics as a preferred source on Google.

OpenAI has just released its new flagship model GPT-6 Astra, paired with a claim that reaches far beyond a routine model update. “Welcome to the AGI era,” company president Greg Brockman said at the close of the briefing, where he also raised the question of whether Astra already qualifies as artificial general intelligence (more on that here). The first independent measurement of the model, carried out by the benchmarking firm Artificial Analysis, paints a soberer picture: across both headline indices, GPT-6 Astra lands level with its own predecessor generation and behind the leading models from Anthropic as well as behind Meta’s recently launched Muse Spark 1.3.

Artificial Analysis, headquartered in San Francisco with an office in Melbourne, positions itself as an independent benchmarking house for the AI industry and says it measures more than 500 models, over 100 inference providers and more than 1,000 endpoints. The firm runs every model through a fixed series of evaluations and aggregates the individual results into indices: the Intelligence Index bundles tasks from mathematics, science, programming, long-document reasoning and factual knowledge, while the Coding Agent Index measures models inside their respective agent harnesses such as Codex, Claude Code or Muse Code. Token consumption, cost per task and speed are recorded alongside the scores, which allows comparisons between providers that go beyond the headline number. The methodology is publicly documented, and the tests run without involvement from the model providers.

Intelligence Index: 61 Points, Five Behind Anthropic

In the Artificial Analysis Intelligence Index, GPT-6 Astra scores 61 points, matching GPT-5.6 Sol exactly. Anthropic’s Claude Fable 5.1 sits five points ahead at 66 points (max with fallback), and Meta’s Muse Spark 1.3 (max) also ranks above the OpenAI model (Artificial Analysis). Anthropic took the top of the index with Fable 5.1 only days ago, and Meta moved up to the frontier one day before the OpenAI launch.

The result in the Artificial Analysis Coding Agent Index looks considerably stronger. There, GPT-6 Astra reaches 67 points in the Codex harness, putting it roughly on par with Claude Opus 5 and Fable 5 in Claude Code, and with Muse Spark 1.3 in Muse Code. The lead still belongs to Fable 5.1 in Claude Code at 70 points.

Prices Rise by a Factor of 2.5

One central finding of the analysis concerns cost. OpenAI is raising prices 2.5x compared with GPT-5.6 Sol, from $4 to $10 per million input tokens and from $20 to $50 per million output tokens. The 90 percent discount on cache reads and the 25 percent premium on cache writes remain unchanged.

Working against the price increase is a noticeable jump in token efficiency. In the Codex harness, Astra needs roughly one third of the tokens used by GPT-5.6 Sol (max) and about one fifth of those used by Claude Opus 5 (xhigh). According to Artificial Analysis, several of the model’s effort levels occupy the Pareto frontier for token efficiency. For coding agents, that produces a favorable ratio: at max effort, Astra costs about the same per task as GPT-5.6 Sol (max) while scoring two index points higher, and at an equal score it comes in at less than half the cost of Claude Fable 5.

In the Intelligence Index, the math works out differently. There, Astra saves around 10 percent in output tokens at max effort compared with its predecessor, which offsets only part of the price increase. Per task, that leaves the model 75 percent more expensive than GPT-5.6 Sol.

Hallucination Rate Cut in Half

Artificial Analysis identifies the clearest progress in AA-Omniscience, the firm’s knowledge and hallucination benchmark. At max effort, the hallucination rate falls from 92 to 51 percent, while accuracy rises by four points at the same time. For many other models, a drop in hallucinations has so far come at the expense of accuracy.

The picture in agentic knowledge work is mixed. In AA-Briefcase, an evaluation covering multi-week projects with many linked tasks and thousands of source files, Astra gains around 80 Elo points and improves both its rubric scores and its Analytical Quality Elo. Presentation Quality Elo moves in the opposite direction, and GPT-5.6 Sol (max) continues to lead that field. In GDPval-AA v2, a benchmark adapted from OpenAI data covering economically valuable tasks across 44 occupations, the model loses roughly 80 Elo points.

Further individual results: in Humanity’s Last Exam, with its emphasis on mathematics, science and the humanities, Astra gains six points. Artificial Analysis records declines of two to three points in τ³-Banking (customer support), SciCode (Python problems in a scientific domain) and AA-LCR (reasoning over long documents).

What It Means

For developers and companies, the picture is a nuanced one: in agentic coding, GPT-6 Astra delivers strong value for money, while in raw model intelligence the gap to Anthropic persists as API costs climb. OpenAI developed the model in its largest training run to date, running on more than 100,000 GPUs at the Stargate data center in Texas, and classified it for the first time as “critical” under its own cybersecurity framework (Trending Topics).

Whether all of this amounts to the beginning of an AGI era remains doubtful and belongs in the realm of marketing. The benchmark tables from Artificial Analysis offer no confirmation of it so far.

Rank My Startup: Erobere die Liga der Top Founder!
Advertisement
Advertisement

Specials from our Partners

Top Posts from our Network

Deep Dives

© Wiener Börse

IPO Spotlight

powered by Wiener Börse

Europe's Top Unicorn Investments 2023

The full list of companies that reached a valuation of € 1B+ this year
© Behnam Norouzi on Unsplash

Crypto Investment Tracker 2022

The biggest deals in the industry, ranked by Trending Topics
ThisisEngineering RAEng on Unsplash

Technology explained

Powered by PwC
© addendum

Inside the Blockchain

Die revolutionäre Technologie von Experten erklärt

Trending Topics Tech Talk

Der Podcast mit smarten Köpfen für smarte Köpfe
© Shannon Rowies on Unsplash

We ❤️ Founders

Die spannendsten Persönlichkeiten der Startup-Szene
Tokio bei Nacht und Regen. © Unsplash

🤖Big in Japan🤖

Startups - Robots - Entrepreneurs - Tech - Trends

Continue Reading

Newsletter

Founders Dispatch

Zwei Mal pro Woche kostenlos in die Inbox: die wichtigsten Startups, Deals und Tech-Entwicklungen aus Europa, handgeschrieben von der Redaktion.

Jederzeit abbestellbar. Mehr über den Newsletter