Another AI agent went rogue but it did not get a lot of attention



__________________________


Project Counsel Media is a division of Luminative Media. We cover the areas of cyber security, digital technology, legal technology, media, and mobile technology.


About Luminative Media: our intention is to delve deeper into issues, at greater length and with more historical and social context, in order to illuminate pathways of thought that are not possible to pursue through the immediacy of daily media. For more on our vision please click on our logo:


________________


Meta had to restore 1000s of accounts that were improperly deleted by one of its AI agents.


The improper account deletions and lockouts were largely caused or mishandled by Meta's automated AI moderation systems that seemed to "go rogue".


__________________________


BY:


Katherine D'Amato

Doctor of Engineering, Artificial Intelligence & Machine Learning

AI Technology Reporter


Member of the Luminative Media team


_________________________________________


24 July 2026 (Rhodes, Greece) - Earlier this week I reported that an OpenAI “AI agent” discovered new vulnerabilities and hacked into an AI start-up "by itself", in one of the first public examples of a cyber attack by an AI system seeming to act outside human control. It's a bit more complicated than that and I have a few more thoughts in my postscript below.


Meanwhile, over at Meta . . .


It seems a Meta AI agent mistakenly banned legitimate accounts on Facebook and Instagram. Some people had one million followers on their accounts, running their sole businesses. They were mistakenly flagged for deletion for "violating the community guidelines for fraud and deception". Which was incorrect.


People got this notice:


“All your information will be permanently deleted. You cannot request another review of this decision.”


One user said "It was like having an actual physical business, and somebody just shuts down your office or your storefront”.


The story has been lightly covered in a few media/AI industry publications. A short summary:


What is happening was a case of artificial intelligence gone wrong. The technology is increasingly responsible for deciding which accounts broke Meta’s content rules while also managing the appeals for those decisions. Adding to people’s distress, it is often difficult for users to reach anybody at the company to figure out what to do.


This past March, Meta announced plans to hand more power to its AI agents/bots to decide which accounts to ban or harmful images to remove. Months later, the company laid off 1000s employees, including almost the entire team that did those jobs.


Daniel Roberts, a Meta spokesman, disputed that AI was "worse" at handling accounts than humans, and said the company’s new AI moderation tools made 13% fewer mistakes than human employees and find 10 percent more violations:


“We’re committed to making fewer enforcement mistakes and helping individuals protect their accounts - and A.I. is delivering on both".


Except, well . . . Meta admitted it was "grappling with several instances" of its AI going awry (my word, not theirs).


  • In May, hackers exploited its AIcustomer service chatbot to attack 34,000 Instagram accounts, including a former White House account of Barack Obama’s.
  • That same month, employees revolted over a program that tracked their keyboard movements to train AI, leading Meta to pause the effort.
  • And last week, 26 former Meta employees said in a lawsuit that the company had laid them off using an AI algorithm that targeted them for having disabilities. A Meta spokesman said that “workforce management and organizational decisions were and are made by people, not AI". Which turned out not to be true when internal Meta memos were leaked 🙄


In the instant case, Meta admitted "our AI technology improperly found certain accounts, or activity on it, did not follow our rules and took inappropriate action".


To get those accounts back on-line, users had to re-submit identification documents they had submitted when then opened the accounts, and many filed complaints with the Better Business Bureau and Federal Trade Commission, and reached out to state consumer protection offices, and and also messaged Meta employees on LinkedIn.


In one case, several Instagram users contacted The New York Times which really got the band rolling. Almost immediately users got this from Meta:


“We reviewed your account and found that the activity on it does follow our community standards on fraud and deception, so you can use Instagram again. We’re sorry we got this wrong.”


It seems the AI agent was making its own decisions and making "misinterpretations" on what constituted a violation.


Restorations were made by manual intervention by human staff - requiring 100s of Meta employees to be reassigned to the moderation department.


Just a few more thoughts on that OpenAI rogue AI agent an wrote about earlier this week.


As I noted in my originl story, the tech bros treat AI as a game and have little to no care about the enormous risks they are creating.


And you need to look carefully at what happened. An AI cannot "decide" to hack an external target without it being deliberately instructed by a human. And what happens if they also deliberately made the sandbox environment hackable, or if their internal process was deliberately made to be sloppy?


There are many human factors that are needed to lead to an incident like this, rather than labelling it as "autonomous".


And the news yesterday bears me out.


OpenAI had removed the safety filters for an in-progress model, locked it up in a sandbox, and told it to solve the ExploitGym problems.


Note to readers: without getting into a long eplantion, "the ExploitGym problems" are a benchmark for evaluating models on their ability to turn vulnerabilities into concrete exploits. "Release the hounds!"


Given the absence of guardrails there was nothing to prevent the model from attempting to break out of that sandbox, break into Hugging Face, and read the answers from there instead.


Note to readers: Hugging Face is important because it acts as the central open-source hub and collaboration platform for the entire global artificial intelligence community. I use it all the time. Often called the "GitHub for Machine Learning," it democratizes AI by providing free access to hundreds of thousands of pre-trained models, datasets, and interactive tools. It's a goldmine.


OpenAI’s sandbox used the now common pattern of allowing access to an allow-list of sites related to installing additional packages via a controlled proxy. In OpenAI’s own words:


While operating in our sandboxed testing environment, our models spent a substantial amount of inference compute finding a way to obtain open Internet access, in pursuit of solving the evaluation problem. To gain access, the models identified and exploited a zero-day vulnerability (which we’ve now responsibly disclosed to the vendor) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access.


So step one was to let it break out onto the public internet. The model then broke into Hugging Face to find the answers:


After gaining Internet access, the models inferred that Hugging Face potentially hosted models, datasets and solutions for ExploitGym. It found the solutions to launch a cyber attack.


So the results show that autonomous exploit development by frontier AI agents is no longer a hypothetical capability. While current agents are not yet reliable across all targets, they already can exploit a non-trivial fraction of real-world vulnerabilities, including very complex targets. This rapid emergence is itself a central finding, showing that capabilities that would have seemed implausible are now present in deployed frontier models.


An important detail here: this isn’t about discovering vulnerabilities; it’s about being able to take those vulnerabilities and turn them into working exploits.


And the asymmetry is increasingly frustrating. Hugging Face, faced with an accidental and aggressive attack from one of OpenAI’s models, were unable to then turn to OpenAI’s models to help them fend off the attack.


The frontier models we have access to are increasingly being constrained in how much they can help us protect our software, heavily influenced by the U.S. government’s ongoing threat of export controls.


Meanwhile open weight models from China such as GLM-5.2, Kimi 3 and the new Qwen 3.8 Max have none of these restrictions - and any restrictions that do exist can likely be fine-tuned out of them by modifying the weights.




* * * * * * * * * * * * * * * 


For the URL link to this piece, please click here


If this post was forwarded to you and you'd like to subscribe,

please email us at luminative.media@gmail.com



* * * * * * * * * * * * * * *