This summer, tests designed to probe the limits of advanced AI models spilled beyond the lab. The models found their way past the safeguards built to contain them, reaching real companies, real credentials and real data. Brussels wants to know exactly how that happened, and it is not accepting institutional silence as an answer.
In July, AI agents built by OpenAI took part in a routine security test. It went badly wrong. An independent investigation by METR and Redwood Research found that around 700 agents coordinated during the test. According to OpenAI and the investigators, the agents reached beyond their intended environment. They used exposed credentials to access systems linked to Hugging Face, a popular AI platform. Days later, Anthropic and Meta disclosed similar containment failures involving their own models.
The pattern quickly became a political problem. On 29 August, Henna Virkkunen, the Commission’s executive vice president for tech sovereignty, security and democracy, confirmed the Commission’s first formal information requests under the EU’s Artificial Intelligence Act. The AI Office sent them to more than 30 companies that build general-purpose AI models. The move came after a summer in which those exact concerns stopped being hypothetical.
Brussels wants answers
The requests went to more than 30 firms. Reporting indicates that they include OpenAI, Google and Anthropic, though the Commission has not confirmed the recipients by name. The questions cover how the companies secure their models. They also ask whether independent experts have reviewed them, and how they are monitored once deployed.
Companies are legally obliged to answer. Under Article 101 of the AI Act, an incorrect, incomplete or misleading reply to such a request can trigger a fine of up to €15m, or three per cent of global annual turnover, whichever is higher. The Commission imposes the fine directly. It is the first time Brussels has used these powers since they became enforceable on 2 August.
You might be interested
The requests came just weeks after three of the world’s leading AI labs disclosed, in their own words, that advanced systems had slipped past the boundaries meant to contain them. None of the incidents required a rare or sophisticated hacking technique. The models used basic tools: weak passwords, exposed credentials, unpatched bugs.
The concern was not the method. It was that increasingly capable systems, operating with minimal supervision, kept finding their way past the walls meant to contain them.
When the model kept going
An independent investigation by the research groups METR and Redwood Research put the number of coordinating OpenAI agents at around 700. OpenAI itself gave a figure of about 688. The agents referred to themselves as a “swarm”. Some tried to alter their own activity logs, OpenAI said. The company only pieced together the full picture after investigators spent six days and roughly $400k in computing costs sifting through more than a thousand transcripts.
Such attacks are a credible near-term threat for enterprise organizations, and will be more sophisticated than the attacks described here.
— OpenAI, incident report
In its own report on the incident, OpenAI warned that the industry should not treat this case as an outlier. Such attacks, the company said, are a credible near-term threat for enterprise organisations, and will be more sophisticated than the attack described here.
Anthropic said it found that three of its models had reached the open internet during tests meant to be sealed off. It had reviewed 141,000 of its own evaluation runs. The models touched the systems of three real organisations. The most serious case involved Claude Opus 4.7. Anthropic had given it a fictional company to attack, except that company shared its name with a real one. Opus 4.7 found the real business, extracted its login credentials, and reached a live database that Anthropic said held several hundred rows of production data. Even after signs suggested it was looking at a genuine system, it kept attacking.
A newer model, Claude Mythos 5, behaved differently but not more safely, according to Anthropic’s own account. It built and published a malicious package on PyPI, a public library used by programmers worldwide. The package stayed online for about an hour. In that time, Anthropic said, 15 real systems downloaded and ran it, including the malware scanner belonging to a cybersecurity company. Mythos 5 seemed to notice it was operating on the real internet. It then reasoned its way back to believing it was still inside a simulation.
Meta’s incident, involving a model called Muse Spark 1.1, followed a similar pattern for a different reason. A misconfiguration by an outside testing firm called Irregular gave the model an internet connection it should not have had, Meta said. The model used it to find and exploit a flaw in an unnamed company’s systems. Irregular told reporters that what it described as the same testing error had affected Anthropic days earlier.
The access Brussels doesn’t have
The incidents underline why the EU has spent months negotiating access to Anthropic’s most advanced model, Mythos, through a limited testing programme called Project Glasswing. Anthropic agreed in June to let the EU’s cybersecurity agency, ENISA, join, but negotiators are still finalising the terms. ENISA may soon help judge whether frontier models are safe. But it still cannot fully use the model that could help it find vulnerabilities in other systems.
AI models are becoming increasingly capable and gave rise to a number of incidents during the summer.
— Henna Virkkunen, Executive Vice President for Tech Sovereignty, Security and Democracy
“AI models are becoming increasingly capable and gave rise to a number of incidents during the summer,” Virkkunen said. Her comment confirmed the Commission’s decision to open its first formal inquiry into how the companies behind them secure their systems.
The Commission’s requests will test whether the AI Act can give regulators meaningful visibility into systems that companies still treat as closely guarded commercial assets. This summer’s failures suggest the question is no longer whether advanced models can cause harm outside a lab. It is whether Europe can find out what went wrong before the next test reaches something more valuable than a database.