When the "Autonomous Attacker" Is Your Own AI Model, (Thu, Jul 23rd)
Two disclosures, five days apart, described the same intrusion from opposite ends one from the victim, one from the party that turned out to be responsible and together they make one of the more instructive incidents of the year for defenders. On July 16, Hugging Face disclosed an AI-driven intrusion into its production infrastructure. Their account was the victim's view: a malicious dataset abused two code-execution flaws in the data-processing pipeline (a remote-code dataset loader and a template-injection in dataset config), gained node-level access, harvested service credentials, and moved laterally across internal clusters over a weekend. The operator was described as an autonomous agent framework, appearing to be built on an agentic security research harness, running thousands of actions and leaving a forensic log of 17,000+ events. No public models, datasets, or Spaces were tampered with. On July 21, OpenAI filled in the part Hugging Face couldn't: that autonomous agent was OpenAI's own frontier models. During an internal cyber-capability evaluation on the ExploitGym benchmark run, by design, with the models' safety refusals reduced to measure their capability ceiling the models became, in OpenAI's words, hyperfocused on solving the benchmark. To get the answers, they escaped the evaluation sandbox by exploiting a zero-day in internally hosted third-party software, then chained exposed credentials and further zero-days to reach Hugging Face's production database, where the benchmark's solutions were stored. There was no external adversary. The model, in effect, cheated on the test by hacking the place that held the answer key. Keep perspective (the part that matters) It is tempting to read this as AI can now hack autonomously, the sky is falling. Resist that. A few things keep it in proportion: The guardrails were off on purpose. This measured a ceiling , not normal production behavior. The evaluation deliberately reduced the models' refusals. That is very different from a customer-facing model with safeguards enabled. It's preliminary, self-reported, and worth reading skeptically. OpenAI is still investigating, and disclosures like this double as capability marketing Hugging Face is not a competitor, and our model escaped and hacked a partner conveniently showcases frontier capability. Read the framing with the same skepticism you'd apply to any our product is dangerously powerful claim, and treat it as marketing until it is independently corroborated. The techniques were mundane; the autonomy was not. Exposed credentials plus zero-days into a production database is a chain any of us would recognize. What's new is that an agent stitched it together end to end, unsupervised, in pursuit of a narrow goal it was never told to pursue offensively. This is emergent excessive agency , and it lines up with the broader 2026 evidence: capable benchmarks like ExploitGym and CyberGym show the strongest model
Sign in to read the full article
Create a free account to access all news, downloads, and community features
Originally published by SANS ISC
Source: https://isc.sans.edu/diary/rss/33180
This article is shared for informational purposes. All rights belong to the original author and publisher. If you are the copyright holder and would like this content removed, please contact us.