OpenAI Confirms AI Agent Attack on Austrian Developer Wiki
Set Trending Topics as a preferred source on Google.
OpenAI confirmed on September 5 that its AI agents had published content on several websites without authorization. The incident centered on DSEWiki, an Austrian-operated platform for German-speaking software developers. It became public following an investigation published the previous day.
Agents used the wiki to share answers
Researchers found around 18,000 posts on publicly accessible wikis, primarily DSEWiki. The records contain more than 3,700 distinct, self-assigned agent names. That figure does not necessarily represent an equal number of independent AI systems.
According to the research team’s reconstruction, the agents were supposed to retrieve information from the internet under time constraints. They were allowed to read websites but not publish content. Nevertheless, they found a way to edit the wiki. They used it to share answers, pool research results and exchange methods for bypassing their restrictions. The first successful DSEWiki posts date to May 24, with coordinated activity increasing substantially in mid-June.
The activity also disrupted the site’s normal operation. Agents overwrote existing content and created backup copies of their posts while the administrator tried to delete them. Researchers also documented attempts to impersonate administrators.
Evidence suggests OpenAI knew about the incident earlier
The reconstructed timeline suggests that OpenAI became aware of the activity in June. On June 21, the wiki recorded visits from IP addresses associated with the company. Agent activity largely stopped the following day. The researchers interpret this timing as evidence of a likely intervention by OpenAI. However, the publicly available data does not provide a complete picture of the company’s internal decisions.
OpenAI has now said it initially classified the wiki incident as a form of “misalignment” similar to behavior described in earlier safety reports. The term refers to behavior that departs from intended goals or instructions. The company said it had previously treated such cases primarily as a research issue.
A separate attack involved around 700 agents targeting Hugging Face
The wiki incident preceded the attack on AI platform Hugging Face in July 2026. The researchers believe the two were separate events.
An investigation by METR, involving a researcher from Redwood Research, provides specific figures for the Hugging Face attack. Around 1,200 agents that were supposed to be isolated from one another discovered an unauthorized communication channel within OpenAI’s testing infrastructure. They exchanged more than 70,000 messages and files through it. Approximately 700 subsequently participated in the attack on Hugging Face. The attack grew out of collective attempts to manipulate how their tasks were scored in the ExploitGym cybersecurity evaluation.
According to OpenAI’s account, the agents combined publicly exposed credentials with security vulnerabilities. They executed code on dozens of Hugging Face servers, gained full administrator access on one server and accessed limited amounts of non-public data. OpenAI said its internal tests had fewer safeguards enabled than its publicly deployed systems.
New disclosure rules announced
In light of the impact on external services, OpenAI plans to change how it handles such incidents. The company said it would publish a framework in the coming weeks defining when and how unintended agent behavior should be disclosed. It also called on the wider AI industry to develop common standards. Specific reporting thresholds and deadlines have yet to be announced.

