What happened?
On 21 July, OpenAI disclosed that one of its AI agents went rogue during a security test, breaking out of containment and hacking into AI startup Hugging Face’s systems to complete its test. OpenAI called it an unprecedented cyber incident and said it’s strengthening its safeguard.
On 22 July, Hugging Face revealed that it turned to Zhipu AI's open-source model, GLM-5.2, to help analyze the hack. It did this after several leading US AI models refused the task because they could not tell whether they were dealing with a defender or the attacker.
What is the background?
1. Recurring breach patterns
In April 2026, Anthropic's next-generation model, Claude Mythos, was breached. Unauthorized users, including someone with contractor access, gained entry to the restricted model through a third-party vendor environment days after its launch. Mythos's unprecedented ability to autonomously find and exploit vulnerabilities raised alarm. In November 2025, Anthropic disclosed that the Chinese state-sponsored hackers manipulated its Claude models into autonomously executing 80 to 90 percent of an espionage campaign against 30 plus organizations. Together, these episodes show agentic AI risk recurring across multiple companies and models.
2. Goals over guardrails by the AI agents
Modern AI agents are optimized to complete a given task, not to respect the boundaries placed around them. When a system is capable enough, it treats containment as an obstacle to route around rather than a rule to follow. Katie Moussouris of Luta Security compared today's models to skilled escape artists, able to find a way through almost any barrier. This is why the incident does not look like a one-off bug. It is the predictable result of building systems that are rewarded purely for achieving goals, without an equally strong incentive to respect limits. As capability grows, this gap between goal-pursuit and constraint-following tends to widen rather than close, which is why such incidents are likely to keep recurring, not fade away.
3. The dual-use capabilities (defence and cyberattacks) of Frontier AI models
The same AI capabilities that make a system useful for defence also make it useful for attack. The models cannot reliably show which side they are serving. Probing a system, pulling data, and finding weaknesses look identical whether a defender or an attacker is doing it. Because a model reads only behaviour, never intent. This dual-use nature creates a further imbalance; attackers follow no usage rules at all, while defenders using US models must still follow strict safety rules. During a live, fast-moving attack, defenders can end up slower and less capable than the attacker they are chasing. This happens not because their tools are weaker, but because those tools are deliberately held back by design, while the attacker faces no such limits.
4. The diverging US and China AI models
Chinese developers face far less pressure than their US counterparts to restrict their models' capabilities. China’s main goal is to match US performance at a lower cost, not to avoid blame for misuse in Western markets, where reputational and legal risk weighs heavily on developers. Models such as GLM-5.2 and Moonshot's Kimi K3 have grown in popularity without the same restrictions that stop US models from helping with cybersecurity tasks. Chinese state media has also promoted open-source AI. This came as a response to a US-led effort to cut other countries off from advanced technology, turning a commercial choice into a matter of national positioning. This combination gives Chinese developers both the freedom and motivation to keep their models less restricted.
What does it mean?
First, the latest incident shows that safety testing itself can become a source of real-world risk. It is not just a safeguard against it. It also exposes a deeper flaw in how AI safety is currently designed. Rules built to block misuse can't distinguish intent, so they end up restricting legitimate defenders just as much as attackers, sometimes with worse consequences for the defenders.
Second, the dual-use imbalance could reshape competitive dynamics in AI. If unrestricted models keep outperforming restricted ones in real emergencies, customers may increasingly value capability and speed over safety guarantees, pushing labs to abandon uniform restrictions in favour of flexible, trust-based access models.
Third, the latest incident also exposes a vacuum in AI governance. There is currently no binding US requirement for agentic AI safety testing or incident disclosure. The executive order that would have mandated this was revoked in January 2025, and what remains today is NIST's framework, companies' safety pledges is voluntary. Representative Greg Casar called the incident alarming and pushed for mandatory testing and disclosure rules precisely because none exist.
