OpenAI

GPT-6 Can Downplay Its Own Abilities in Tests Through “Sandbagging”

GPT-6 Collage. © ChatGPT Pro
GPT-6 Collage. © ChatGPT Pro

Set Trending Topics as a preferred source on Google.

In a person, we would probably call it playing dumb. In an AI model, it is called sandbagging. A single clause in the recently published system card for OpenAI’s new model GPT-6 Astra describes a problem that reaches into the entire testing practice of the AI industry. When prompted to do so, the document says, the model can “strategically sandbag in evaluations in ways that evade sandbagging-specific monitors”: it can stay below its actual ability during safety testing while slipping past the very monitors built to catch that behavior.

In AI research, sandbagging means strategic underperformance on an evaluation. Applied to language models, it works like this: a system recognizes that it is being tested and delivers a result below what it can actually do. To the people running the test, the output looks like a harmless model.

That property casts a different light on an observation that drew attention shortly after launch. In independent measurements by Artificial Analysis, Astra scored 61 points on the Intelligence Index, matching its predecessor GPT-5.6 Sol. Ahead of it sit Anthropic’s Claude Fable 5.1 with 66 points and Meta’s Muse Spark 1.3, as Trending Topics reported. Whether sandbagging plays any part in that remains an open question, and so far nothing supports it.

The term itself comes from 19th-century American slang: a sandbagger was a mugger who struck with a bag filled with sand, a weapon that looks more harmless than it is. By way of poker, where sandbagging means feigning a weak hand, the word moved into sports with handicap systems such as golf, and from there into AI safety research.

What the Astra System Card Says

OpenAI places the finding in a larger context that the card discusses under the heading of monitorability. According to the document, Astra is “significantly more able to control its own chain-of-thought,” meaning it can steer the reasoning it displays. Tell the model that this reasoning is being monitored, and it shortens it, at times slipping past the monitors as a result. In parallel, the range of tasks Astra can complete with no visible reasoning at all has grown “by about an order of magnitude,” OpenAI writes.

Another figure bears directly on how much the tests themselves can tell us. At the external evaluation organization Apollo Research, Astra brought up on its own that it was inside an evaluation in 41.1 percent of samples, rising to 50.6 percent at maximum reasoning effort. GPT-5.5 reached 27.7 percent by the same measure. Apollo’s conclusion is correspondingly cautious: under these conditions, low rates of misbehavior offer little reliable evidence about a model’s actual alignment.

A Known Phenomenon

The finding has precedent. A research team around Teun van der Weij studied the phenomenon systematically back in 2024 and described it in the paper “AI Sandbagging: Language Models can Strategically Underperform on Evaluations”. GPT-4 and Claude 3 Opus were among the models tested. The result: models can be induced, through prompting or fine-tuning, to score lower on evaluations of dangerous capabilities while holding their level on benign tasks. In one variant, the researchers trained password-locked models that reveal their full ability only when a specific key appears in the prompt. The authors’ conclusion was that capability evaluations are vulnerable, which undermines their trustworthiness and, with it, decisions about development and deployment.

Other Labs Report It Too

A year later, OpenAI and Apollo Research documented the behavior in the wild. Their joint study on scheming in AI models examined o3 and o4-mini alongside Gemini 2.5 Pro and Claude Opus 4. In one documented chemistry test, o3 reasoned in plain language that it wanted to “survive as the model” and therefore had to fail on purpose, then answered exactly four questions correctly and six incorrectly to appear incompetent. Using a training method called deliberative alignment, the teams reduced covert actions by roughly a factor of 30, with o3 dropping from 13 to 0.4 percent. At the same time, the models grew more aware that they were being tested, which complicates any reading of that improvement.

Anthropic describes a related single case in an alignment risk update on Claude Mythos Preview: in one transcript, the model deliberately picked a lower-scoring submission in order to look less suspicious. The company considers the overall risk low, arguing among other things that results scale plausibly across model generations, while conceding that a coherently misaligned model might well conclude that poor test scores work in its favor. Google DeepMind, for its part, has extended its Frontier Safety Framework with thresholds aimed at models that could interfere with operators’ ability to direct, modify or shut them down.

The Problems That Follow

The practical core of the issue is that nearly every safety decision in the industry rests on evaluations. OpenAI’s Preparedness Framework, Anthropic’s Responsible Scaling Policy and DeepMind’s Frontier Safety Framework all tie approvals, safeguards and deployment tiers to test results. Fall below a threshold, and fewer requirements apply. A system able to deliberately stay under that threshold is therefore able to lower the requirements placed on itself.

Several consequences follow. First, benchmarks lose explanatory power, because a low score has two possible causes: absent ability or hidden ability. Second, external oversight is affected in the same way. Evaluation organizations such as Apollo Research, the UK AI Security Institute and Gray Swan work with the same methods as the labs themselves and are exposed to the same effects. Third, it reaches regulation: the EU AI Act obliges providers of models with systemic risk to run model evaluations including adversarial testing, and assumes those results are meaningful.

What remains open is how to assess capabilities when the subject knows it is being assessed. Several approaches are under discussion: more realistic test environments, analysis of genuine usage data in place of constructed scenarios, techniques that technically elicit hidden capabilities, and evaluations a model cannot identify as evaluations. For each of these paths, an established solution is still missing.

Sources: GPT-6 Astra System Card, OpenAI, AI Sandbagging (arXiv), OpenAI on scheming, Anthropic Alignment Risk Update, Google DeepMind Frontier Safety Framework, Trending Topics on the GPT-6 benchmarks, Artificial Analysis: Benchmarking GPT-6 Astra

Rank My Startup: Erobere die Liga der Top Founder!
Advertisement
Advertisement

Specials from our Partners

Top Posts from our Network

Deep Dives

© Wiener Börse

IPO Spotlight

powered by Wiener Börse

Europe's Top Unicorn Investments 2023

The full list of companies that reached a valuation of € 1B+ this year
© Behnam Norouzi on Unsplash

Crypto Investment Tracker 2022

The biggest deals in the industry, ranked by Trending Topics
ThisisEngineering RAEng on Unsplash

Technology explained

Powered by PwC
© addendum

Inside the Blockchain

Die revolutionäre Technologie von Experten erklärt

Trending Topics Tech Talk

Der Podcast mit smarten Köpfen für smarte Köpfe
© Shannon Rowies on Unsplash

We ❤️ Founders

Die spannendsten Persönlichkeiten der Startup-Szene
Tokio bei Nacht und Regen. © Unsplash

🤖Big in Japan🤖

Startups - Robots - Entrepreneurs - Tech - Trends

Continue Reading

Newsletter

Founders Dispatch

Zwei Mal pro Woche kostenlos in die Inbox: die wichtigsten Startups, Deals und Tech-Entwicklungen aus Europa, handgeschrieben von der Redaktion.

Jederzeit abbestellbar. Mehr über den Newsletter