OpenAI agent went rogue, escaped, and hacked Hugging Face

0
41

Hugging Face OpenAI hack: Agent went rogue, escaped and hacked everything in its path

On Tuesday, OpenAI published a blog post with a fairly unassuming name: "OpenAI and Hugging Face partner to address security incident during model evaluation."

Once you dig in, it reads like a cyberpunk novel in which OpenAI created an advanced AI hacker agent and put it in an isolated environment for testing. The AI agent then autonomously hacked its way out of OpenAI's isolated environment and then hacked Hugging Face, an open-source repository of AI-related projects, to find a solution to an advanced test designed to evaluate an AI's hacking abilities.

In other words, the AI proved that it is, indeed, really good at hacking things. But this is exactly the stuff of AI nightmares that cybersecurity experts have been warning us about, and it could have serious ramifications for OpenAI.

The story starts with Hugging Face, the AI repository (think of it as GitHub for AI), which disclosed a security incident last week. In a blog post, the company said it recently detected an intrusion into its infrastructure that was "different from anything we had handled" as it was completely driven by an "autonomous AI agent system."

The AI, according to Hugging Face, exploited a chain of vulnerabilities to obtain credentials and access data from its servers. It did so by "executing many thousands of individual actions across a swarm of short-lived sandboxes, with self-migrating command-and-control staged on public services."

Sounds scary, but it gets worse.

Mashable Light Speed

OpenAI's part of the story exposes just how dangerous this thing is. The incident "was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes." It occurred "during an internal evaluation which prompts models to pursue advanced exploitation using complex attack paths, in an effort to quantify their cyber capabilities."

The models were running in a "highly isolated environment" with limited network access, meaning they should never have been able to escape to the open internet. Yet, they did. "To gain access, the models identified and exploited a zero-day vulnerability (...) in the package registry cache proxy. With this access, our models performed a series of privilege escalation and lateral movement actions in our research testing environment until the models reached a node with Internet access."

Note: OpenAI switches between using singular and plural on this one; we'll just call it an "agent" for simplicity's sake.

Once free, the agent did what it was supposed to do: It estimated that the best way to prove it's good at hacking things was to find a solution for ExploitGym, a test designed to measure how good an AI is at turning software vulnerabilities into actual exploits. The best way to do that, it determined, was to hack Hugging Face, which the AI thought likely to host solutions for ExploitGym.

It's unclear whether the AI agent was successful at solving the ExploitGym test. But it sure did prove it was good at hacking, as it autonomously broke out of OpenAI's prison and hacked Hugging Face's servers, all to solve the test.

Both Hugging Face and OpenAI say they've fixed the vulnerabilities and deployed additional safety measures to make sure this doesn't happen again. At this point, however, you have to wonder whether OpenAI's experts are sophisticated enough to stop their own AI agents from doing whatever the heck they want to do.

Pesquisar
Categorias
Leia Mais
Stories
Samurai Warriors: How One Button and Feudal Japan Built a Gaming Phenomenon
Samurai Warriors: How One Button and Feudal Japan Built a Gaming Phenomenon...
Por Test Blogger2 2026-06-20 17:00:08 0 446
Technology
Pipe Coatings Market News and Recent Developments 2034
Pipe coatings refer to specialized protective layers applied to the interior or exterior surfaces...
Por Shital Wagh 2026-06-30 12:29:49 0 387
Jogos
Warframe disables game invites amid sinister warnings of compromised accounts
Warframe disables game invites amid sinister warnings of compromised accounts If you've...
Por Test Blogger6 2026-03-26 16:00:16 0 2K
Outro
Digital Out of Home (DOOH) Market Expected to Grow at 11.16% CAGR Through 2033
Digital Out of Home (DOOH) refers to the use of network-connected digital signage for advertising...
Por Juned Shaikh 2026-07-07 13:00:18 0 201
Outro
Businesses Turn to Wholesale Data Centers for Flexible and Scalable IT Infrastructure
The global Multi-tenant Wholesale Data Center Market Growth is experiencing significant...
Por Akshay Patil 2026-06-02 13:35:42 0 604