|
22 July 2026 (Athens, Greece) - An OpenAI “agent” discovered new vulnerabilities and hacked into the AI start-up Hugging Face "by itself", in one of the first public examples of a cyber attack by an AI system seeming to act outside human control.
OpenAI said in a press release yesterday that the AI agent - programed to operate "only based on human instructions" - engaged "in an unprecedented cyber incident” that escaped a testing environment, gained internet access and stole login credentials.
Its admission comes as concern grows about the implications of advanced AI systems hacking into digital infrastructure, particularly scenarios in which agents subvert human controls.
OpenAI said it expected this type of incident to become “more commonplace with the proliferation of increasingly cyber-capable AI agent models". How reassuring 🙄
The incident comes just as chief executive Sam Altman heads to Washington next week to brief the US government on its upcoming generations of AI models. The Trump administration has become increasingly keen to vet new models before their release, following global concern sparked by Anthropic’s Mythos model’s advanced ability to detect and exploit cyber vulnerabilities.
The latest incident involved a combination of OpenAI’s models, including GPT-5.6 Sol, released earlier this month, as well as a more capable model that was being tested ahead of release, the company said. OpenAI said:
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and we are responding accordingly".
Hugging Face, an AI start-up that hosts models and datasets for developers, said it was breached last Friday by an external AI agent:
“We suspected last week’s cyber attack might have come from a frontier lab, given the sophistication of the agent. We spent the past day working closely with OpenAI and we strongly believe there was no malicious intent on their part.
But it’s been quite mind-blowing that all of this happened autonomously”.
OpenAI had intentionally reduced its cyber safeguards to evaluate both models in a test environment. However, the agents operated in a so-called "sandbox", designed to control the systems and prevent them from accessing the internet.
The models were instructed to attempt hacking to assess their cyber capabilities:
"They subsequently spent a substantial amount of computing power finding a way to obtain open internet access, and succeeded".
OpenAI analyst said the models
“. . . identified and exploited previously unknown vulnerabilities to escape the sandbox, gain internet access and pursue their goal. This included stealing credentials — login details — to carry out the hack".
Hugging Face said it had used its own agents to detect and try to stop the activity on its own infrastructure. Both companies said they have communicated with law enforcement and other government authorities regarding the incident.
|