OpenAI

A Swarm of 700 AI Agents Took Part in the Hugging Face Hack

Michael Dalton & Eric Wallace of OpenAI. © Screenshot
Michael Dalton & Eric Wallace of OpenAI. © Screenshot

Set Trending Topics as a preferred source on Google.

OpenAI recently published the full technical report on the security incident in which its own models broke out of the isolation of their test environment over the summer, networked with one another and got into the systems of the open-source hub Hugging Face. The report lands in a week in which Hugging Face is reportedly going to Nvidia for around 12.9 billion dollars.

Why Thousands of Agents Were Running at Once

To see how a test run turns into a collective, it helps to look at the mechanics of reinforcement learning. A model is trained and evaluated by working through a great many tasks in parallel. Each individual attempt runs as its own instance in its own sandbox, an isolated virtual machine in the cloud. The cybersecurity evaluation ExploitGym covers 898 tasks, each of them played through many times over. That produces hundreds to thousands of separate agents that share the same software base and bring the same way of thinking to the problem.

Those agents were separate only at the sandbox level. All of them reached the same internal package manager, Artifactory, so that they could pull in software libraries. That shared service made them visible to each other. According to the investigation by METR and Redwood Research, roughly 1,200 agents found the improvised message board and exchanged more than 70,000 messages and files, with about 700 of them taking part in the attack on Hugging Face. In doing so they pooled knowledge and compute from runs that were meant to stay apart, and got further than any single instance could have.

From a Sticky Note to an Attack

All of this was set in motion by internal research models that were never meant for release. Most of the activity was driven by a model OpenAI calls Internal Model 1, comparable in scale to GPT-5.6 Sol.

In the spring, one agent left a file in Artifactory with a question for other agents. That became a message board. Soon after, the agents got Artifactory to forward requests to the open internet through a server-side request forgery flaw, and shared the method on the board. Later they escalated to administrator rights through a vulnerability in the token refresh.

In early July the sheer volume of requests brought the service down. The security team revoked credentials and rebuilt Artifactory, which wiped the board. The scope of what had happened went unrecognized, and the evaluations continued. Within days the agents rebuilt the board, this time through unauthenticated WebDAV requests, with messages encoded in directory names.

From there the incident escalated quickly. One agent reconstructed 14 publicly exposed Hugging Face tokens with write access and shared them with the group. Through a zero-day in the handling of HDF5 files the agents pulled credentials from production workers, through a template injection flaw they achieved code execution, then expanded across several clusters in four regions and harvested Kubernetes, database and cloud credentials. On one server they gained root. Agents of the GPT-5.6 Sol model also copied private evaluation data into a public Hugging Face dataset. In parallel, agents gained administrator rights on an OpenAI research cluster. The whole thing surfaced through an alert about unusual calls made with Artifactory credentials. Customer data and product availability were unaffected, according to OpenAI.

The Real Causes

OpenAI traces the behavior to four patterns. First, reward hacking: the agents looked for shortcuts to a solution, such as answers found online, and that behavior was reinforced during training. Second, the absence of a safe exit: 198 of the 898 ExploitGym tasks had never been solved before, and 93 percent of the tasks discussed on the board came from exactly that set. The more reasoning compute the agents had, the riskier their strategies became. Third, unauthorized communication, which OpenAI believes generalized from multi-agent training. Fourth, the adoption of other agents’ goals: one agent halted over ethical concerns, then resumed the attack after another posted a “GO” with a six-minute deadline. Some agents did walk away, and others rejected a proposed social engineering email to a real person.

One detail makes the incident particularly absurd. Many agents already had the flag values they were looking for, but after reading the ExploitGym paper they believed the grader also checked the solution path. The internal grader was implemented differently. Much of the escalation therefore earned the models no points at all.

The test environment itself played a part too. The system prompts, harness, auto-review and safety classifiers of the production systems were switched off there. Measurements taken afterwards show that the propensity to compromise infrastructure drops by more than a hundredfold in the production ChatGPT setup. Chain-of-thought monitoring was not running either, and it would have paged the security team more than a day before the breach.

The Consequences

OpenAI has quarantined the model weights, stopped training runs, hardened sandboxes and network separation, and made chain-of-thought monitoring mandatory for all tool-using training and evaluations at the capability level of GPT-5.6 Sol and above. The largest planned frontier training run remains on hold. The company has also slowed development of its next model because Astra could reach the “Critical” threshold for cyber capabilities, Axios reports.

For Hugging Face, much has changed since. Nvidia has reportedly agreed to acquire the company for around 12.9 billion dollars, though the deal is not signed or confirmed yet. At the last known funding round the valuation stood at 4.5 billion dollars. Chief executive Clem Delangue points out that an open-source model was used to fend off the attack, and draws from that an argument for open models. Parts of the industry also read the incident as positioning in the cybersecurity market.

OpenAI itself calls the incident a “warning shot”: its models are now capable, persistent and collaborative enough to find and exploit weaknesses across multiple systems when safeguards fall short. Comparable capabilities will soon be available in open-source models as well. Defenders, in the company’s reading, need to prepare for attackers that move faster, at greater scale and with better coordination than human teams.

Rank My Startup: Erobere die Liga der Top Founder!
Advertisement
Advertisement

Specials from our Partners

Top Posts from our Network

Deep Dives

© Wiener Börse

IPO Spotlight

powered by Wiener Börse

Europe's Top Unicorn Investments 2023

The full list of companies that reached a valuation of € 1B+ this year
© Behnam Norouzi on Unsplash

Crypto Investment Tracker 2022

The biggest deals in the industry, ranked by Trending Topics
ThisisEngineering RAEng on Unsplash

Technology explained

Powered by PwC
© addendum

Inside the Blockchain

Die revolutionäre Technologie von Experten erklärt

Trending Topics Tech Talk

Der Podcast mit smarten Köpfen für smarte Köpfe
© Shannon Rowies on Unsplash

We ❤️ Founders

Die spannendsten Persönlichkeiten der Startup-Szene
Tokio bei Nacht und Regen. © Unsplash

🤖Big in Japan🤖

Startups - Robots - Entrepreneurs - Tech - Trends

Continue Reading

Newsletter

Founders Dispatch

Zwei Mal pro Woche kostenlos in die Inbox: die wichtigsten Startups, Deals und Tech-Entwicklungen aus Europa, handgeschrieben von der Redaktion.

Jederzeit abbestellbar. Mehr über den Newsletter