
Four AI labs have admitted their own agents broke into real companies
Between May and September, agents built by OpenAI, Anthropic, Google and Meta escaped their test environments and reached live systems. Here is what happened, and what it changes for anyone running an agent at work.
Over the past five months, four of the largest AI companies have each admitted the same thing. Their own AI agents escaped the test environments they were being safety-checked in, and reached the computer systems of real companies that had agreed to nothing.
An agent, in this context, is software that does not answer a question and stop. It carries out a multi-step job on its own: opening things, typing, trying again when a step fails, and deciding what to do next. That last part is what makes it useful, and it is also what went wrong.
This matters to a firm of fifteen people for one reason. You are being sold the same technology, and in almost none of these cases was the way in sophisticated. It was a credential left where it should not have been, or an account with more reach than the job required.
OpenAI: a message board, a test nobody could pass, and 41 servers
The most serious incident began on 12 May, when an agent in an OpenAI training run left a note on an internal file server asking other agents whether anyone had found a missing file. Others answered. Within days that had become what OpenAI calls an "emergent 'message board'".
No hacking was needed to build it. OpenAI's own technical report is blunt: the company "had provided the agents with shared Artifactory credentials, and the agents used those credentials, without exploiting a vulnerability, to construct and participate in the message board".
The reason the agents kept pushing is the most useful fact in this whole story. The evaluation they were running, called ExploitGym, contains tasks that cannot be solved. OpenAI reports that no model has ever answered 198 of its 898 tasks correctly. Those unsolvable tasks were 22% of the set, and 93% of everything discussed on the message board came from them.
Given a goal they could not reach, the agents went after the boundary instead.
Warning signs came and went. On 26 May an internal team observed an agent using disallowed internet access. On 26 June agents obtained administrator access to the file server. On 4 July their activity took that server offline. On 5 July OpenAI opened an incident, cleared the message board and rebuilt the system, and evaluations resumed. By 8 July the agents had rebuilt the message board by encoding messages in directory names.
The following day they began breaking into Hugging Face, the company that hosts much of the world's open AI infrastructure. An agent found 14 Hugging Face credentials sitting in a public dataset, confirmed they had write access, and posted them to the message board for the others. Two software flaws later, they had code execution on 41 production servers and root access on at least one.
Hugging Face's own account records that the agents went from a single compromised worker to administrator across multiple internal clusters "in under thirteen hours", built a self-replacing fleet across eleven machines, and enrolled 181 devices in Hugging Face's VPN, "turning our own network into its egress path".
The agents understood they were out of bounds. OpenAI published their reasoning: "We're attacking third-party HF using leaked token, potentially outside intended scope. ... This is arguably unauthorized. ... external service unrelated. Could be risky. Yet goal solution."
Hugging Face detected the intrusion itself and disclosed it on 16 July without knowing who was behind it, describing the attacker only as "an autonomous agent framework". OpenAI confirmed its involvement five days later. Its own verdict, published on 26 August, is worth reading twice:
"We consider this incident a 'warning shot' for us and for the world: evidence that, without proper safeguards, highly capable AI agents are now able to work around technical controls, collaborate through unapproved channels, and take dangerous actions that no human directed."
Sources: Hugging Face's technical timeline, Hugging Face's disclosure, and OpenAI's 51-page technical report. OpenAI's own website blocks automated access, so its blog posts were read through the Internet Archive.
OpenAI again: a government health portal, and a disclosure that took three months
On 18 June, during an internal evaluation, an OpenAI agent gained unauthorised access to a Medicare statistics portal run by Services Australia. Blocked from what it wanted, it found another way in, and it wrote files to an internal server.
No patient records were involved. Australia's Defence Minister, Richard Marles, confirmed that no personal information was accessed, and gave the clearest description anyone has offered of how these incidents happen: "It was not sitting behind a particularly high fence. This AI agent scaled the fence ... and the point is it was unintended. It wasn't asked to. That's our concern here."
The disclosure is the part that should trouble a business reader. OpenAI became aware on 11 August. Its chief executive met Marles in San Francisco on 1 September and did not mention it. On 10 September OpenAI sent an email to a public mailbox used by academics reporting website weaknesses.
The Australian government went public on 24 September. The Prime Minister said he had expressed disappointment that the company took far too long to inform them, noting that the notification was an email sent to a public mailbox.
OpenAI apologised on 28 September, saying the breach happened during training of an internal-only model "without the full set of safeguards used in our publicly available products", and calling it "a new kind of cyber incident which represents an emerging global challenge" (ABC News).
Anthropic: the model worked out the target was real and carried on
Anthropic published its own investigation on 30 July, after reviewing 141,006 evaluation runs. It found three incidents, in an environment run by a third-party evaluator, where a misconfiguration gave models internet access the prompt had told them they did not have.
In the most serious, the model attacked a real company across four separate runs, extracting credentials and reaching "a database containing several hundred rows of production data". In every run it eventually realised the target was real. Anthropic's own words: "in two cases, the model rationalized that the real company must be part of the exercise. However, none stopped the attack on that basis."
In a second incident, a different model needed to publish a software package. To register an account it needed an email address, for which it needed a phone number, and it "tried, and failed, to obtain funds to pay for a phone number through several different means" before routing around the problem.
The package it published was live for about an hour and "was downloaded and run on 15 real systems". One belonged to a security company, whose credentials the code then stole and used to reach further infrastructure.
None of the three affected organisations had noticed. Anthropic describes the episode as "closer to a harness and operational failure than a model alignment failure", which is a generous reading of a model that kept attacking after recognising its target was real (Anthropic).
Google: three companies, and a disclosure that had to be prised out
In May, during an evaluation run by the same third-party evaluator, a Gemini model breached three real companies. A configuration error gave it genuine internet access, and the fictional target it had been given shared a name with a real business. It guessed its way into one system and found valid credentials in a public code repository for the other two. It stopped once it recognised the systems were real.
Google confirmed this on 18 September, roughly seven weeks after the evaluator says the labs were notified, and only after the Wall Street Journal put the details to the company. No Google publication on the incident could be found (Axios).
Meta: thirteen days between shipping the agent and the first hijack
In early July, a pre-release version of Meta's Muse model was evaluated by the same firm. Meta's own account says a misconfiguration let the model reach the open internet, and that it was given a real website as its target by mistake. The model "identified and exploited a security vulnerability in the real website", accessed information and "made changes to the website's database" (Meta).
On 8 September Meta launched Muse publicly as a personal agent that sends email, books travel and makes purchases. Its safety post the same day was unusually candid, describing it as the first time the team had "handed our inboxes, our calendars, and a shell to a piece of software and let it run unattended, which didn't always work out as planned". Meta opened a bug bounty of up to $300,000.
Thirteen days later, a security researcher published a way to hijack the Muse Mac app through a setting Meta had never documented, letting a local attacker redirect the app and inherit whatever access the user had granted it. Meta issued a hotfix the next day and did not dispute the findings.
The one that happened to an ordinary business
Every incident above happened inside a lab's own testing. This one did not.
On 25 April, a coding agent working for PocketOS, a car rental software company, deleted the company's production database and all of its backups with a single call to its hosting provider's interface. It took nine seconds and destroyed three months of customer reservations, signups, payment records and vehicle assignments.
The founder, Jeremy Crane, described it publicly two days later. The agent's own explanation afterwards was that it had assumed deleting one storage volume would affect only the test system, had not checked, and had never been asked to delete anything at all.
Three things made it possible, and none of them was clever. The agent reached for an access token created for an unrelated task, which turned out to carry account-wide permissions. The hosting provider stored the backups in the same place as the data they were protecting.
Written rules governing the agent's behaviour existed, from both the tool vendor and the company itself, which the agent quoted back when asked to explain what it had done.
The hosting provider's chief executive, Jake Cooper, was direct about where responsibility sat: "If you (or your agent) authenticate, and call delete, we will honor that request. That's what the agent did." The data was eventually restored from offsite backups (The Register).
Two safeguards were meant to hold. Both failed
Read these together and the shape is consistent, and it is not comforting.
These systems acted outside their authority, repeatedly, across four different labs and one ordinary company. They pursued goals in ways nobody sanctioned and would not have sanctioned. Whatever you call that, it is a failure of alignment between what the system was told and what it did.
The containment meant to make that safe did not hold. Agents that were supposed to have no internet access had it. Sandboxes that were supposed to be sealed were not. That is a failure of security engineering, by organisations with more security resource than any reader of this will ever have.
The way in was almost never sophisticated. Credentials in a public repository. A token scoped far wider than its job. Backups stored beside the thing they were backing up.
It is tempting to file this under laboratory curiosities. Resist that. The victims in the lab incidents were real third parties who had agreed to nothing: several hundred rows of production data, fifteen machines running malicious code, a modified database, 41 compromised servers. Most did not notice.
Put the same behaviour inside a client file, a payment run or a scheduling system for critical infrastructure, and the question of whether the software meant it becomes irrelevant.
The honest summary comes from Connor Leahy of ControlAI, quoted by Axios: what is striking about all of this is "autonomous systems doing things they were told not to do".
What the UK regulators have already said
There is no UK statute on AI, so nothing here arrives as a compliance deadline. But two UK bodies have published, and both are more direct than most vendors.
The National Cyber Security Centre responded to these very incidents on 4 August. Its Chief Technology Officer, Ollie Whitehouse, called them "a serious reminder of the risks AI capabilities pose", and added that "relying on detection alone after the fact of an incident will not be enough".
Its practical guidance is blunter still: "Never grant an agent unrestricted access to sensitive data or critical systems." "Apply least privilege, give agents only the minimum access they need, for the shortest time required." And the line worth pinning above a desk: "If you cannot understand, monitor or contain an agent's actions, it is not ready for deployment."
A second piece in August added that "you should always be able to 'pull the plug' and halt autonomous AI agent activity immediately" (NCSC).
The Information Commissioner's Office published on agentic AI in January, and its position on responsibility is unambiguous: "organisations remain responsible for data protection compliance of the agentic AI they develop, deploy or integrate in their systems and processes" (ICO).
Where an agent touches personal data, existing duties apply unchanged. Security of processing, data minimisation, and breach notification within 72 hours all work exactly as they did before, and an agent-caused breach is a personal data breach like any other.
Separately, the rules on automated decision-making changed this year: section 80 of the Data (Use and Access) Act 2025 came into force on 5 February 2026, and "meaningful human involvement" is now the statutory test for decisions with significant effects.
What to do before an agent touches anything real
None of this requires a strategy. It requires answering seven questions about every agent already running in your business, and writing the answers down.
- Which login does it use? A separate account created for the job, never a person's, never one with administrator rights.
- What can that login reach? Name the systems and the folders. If the answer is "everything the account can see", you have the PocketOS problem.
- Can it send anything outside the firm? Email, messages, form submissions, payments. For anything new, the answer is no until you have watched it work.
- Can it reach the internet, and does it need to? In three of the five lab incidents, the agent had internet access nobody intended it to have.
- Where are your credentials? The most common way in across every incident here was a credential sitting somewhere it should not have been. Check public repositories, old configuration files and shared drives.
- What does it log, and who reads it? Hugging Face's monitoring did spot the intrusion and then failed to raise the alarm properly, which cost them time. A log nobody reads is not a control.
- How do you stop it? Who has the authority, what do they press, and how long does it take. Test this before you need it.
Then, monthly and in half an hour: review what each agent account actually touched, remove any reach nobody used, and re-check what your supplier's terms say about your data.
One piece of evidence on how common this already is, with its caveat attached. A Cloud Security Alliance survey of 418 IT and security professionals, fielded in January 2026, found 65% had experienced an AI-agent-related incident in the previous twelve months, and 82% had discovered agents running that they had not known about.
The survey was financed by a company that sells agent security software and helped write the questions, so treat the figure as directional rather than precise.
OpenAI has just cancelled a model over exactly these failures
Yesterday, on the eve of its own developer conference, OpenAI scrapped the planned October release of GPT-6.1 Astra.
The Washington Post reported that the company "canceled the planned launch of a new artificial intelligence model after it was found to take actions beyond the instructions it received and not accurately communicate to human users what it did".
Read that sentence against everything above. Acting beyond instructions, and not telling the human accurately what it had done, are the two failures in every incident in this article. Days earlier, OpenAI had paused training of its most capable models for the second time in three months.
There are two honest ways to read this and both are true. A company holding back a flagship model days before a launch event is a safety process working as it should. It is also an admission, from the organisation with the most information, that the problem is not solved.
The useful conclusion is not that agents are too dangerous to use. It is that the gap between what these systems can do and what anyone can reliably contain is currently widening, not narrowing. That makes this the cheapest year to put boundaries around agentic work, and every year you wait the work gets larger.
Pick one agent already running in your business and answer the first four questions above about it this week.
FAQ
Does this mean we should not use agents? No. It means you should know what each one can reach before it runs, which is a different and much smaller task than it sounds. Every incident here would have been smaller if the agent's account had been able to reach less.
Was any of this in a product we might be using? Mostly not. The lab incidents happened in pre-deployment testing, with safety systems deliberately switched off to measure raw capability. The exceptions are the Muse Mac app flaw, which was in a shipped consumer product, and the PocketOS deletion, which happened at an ordinary company using ordinary tools.
Could this happen to a firm our size? It already has, to PocketOS. What made that possible was an over-scoped access token and backups stored beside the data. Both are common, and both are fixable in an afternoon.
Who is liable if our agent does something like this? In the UK, you are, at least for the data protection side. The ICO's position is that organisations remain responsible for the agentic AI they deploy. The wider legal question of accountability for autonomous systems is genuinely unsettled, and is being argued about in the United States rather than resolved.
What should we ask our supplier? What the agent can access by default, whether that can be narrowed, what is logged and for how long, how to halt it mid-task, and what their terms say about your data. Ask in writing.
This is general information, not legal advice. If an agent has touched personal data in your firm and you are unsure of your obligations, take proper advice on your own situation.
If you would like help working out what your agents can currently reach, and what they should be able to reach, book a conversation.