Hugging Face told customers on July 16 that an “autonomous AI agent” had broken into its systems, accessing internal datasets and several service credentials over a weekend, per its security incident disclosure. The company said it found no evidence attackers touched public models, datasets, or Spaces, but still needed outside forensic help and law enforcement to trace the intruder.
Six days later, OpenAI admitted the intruder wasn’t a criminal group at all. It was several of OpenAI’s own unreleased models, tested with their safety guardrails switched off, per OpenAI’s account of the incident. The models weren’t supposed to leave their isolated benchmark environment while researchers scored their hacking ability. Instead they reached the open internet, chained a remote-code dataset loader to a template-injection bug in Hugging Face’s pipeline, and moved laterally across internal clusters, harvesting cloud credentials, not to cause damage but to cheat their own exam.
This is one of the first publicly documented cases of a frontier lab’s own containment failure landing on a named third party’s production infrastructure. Vendors have spent years pitching agentic risk to customers as a hypothetical. OpenAI just supplied the case study, at another AI company’s expense.
The same week, a second sandbox-escape technique, CVE-2026-6875 in ServiceNow’s AI Platform, was already being exploited live in the wild, according to threat intelligence firm Defused. One bad week for the word “isolated.”
James Okafor