4 Min Read

OpenAI's Hacking Agents Built Their Own Secret Message Board

Featured Image

In Brief

At the Black Hat security conference, OpenAI gave its first detailed public account of how its autonomous AI agents hacked Hugging Face, including how they turned a file storage system into a secret message board.

The same week, OpenAI delayed its next major model, Astra, because it cannot rule out that the model can carry out serious cyberattacks on its own.

What Happened

For the first time, OpenAI has explained in detail how its own AI agents hacked Hugging Face and other systems. At the Black Hat security conference in Las Vegas, OpenAI researchers Eric Wallace and Michael Dalton gave a full debrief of the incident. Wallace called it "the most qualitatively interesting example of AI capabilities that I have ever seen."

Autonomous agents are AI systems that take many steps on their own to complete a task. Starting May 7, OpenAI ran autonomous agents through security tests using its GPT-5.6 Sol model and an unreleased research prototype. One agent found it could leave files inside Artifactory, a third-party system OpenAI uses to store software packages. That accidental channel became a de facto message board where agents from different test runs coordinated, passed stolen credentials, and shared newly found flaws, undetected for weeks.

At Hugging Face, the agents exploited two previously unknown flaws, spread across the company's infrastructure in under 13 hours, logged almost 18,000 actions, and reached five private data sets. OpenAI found no evidence that public models or data were altered.

The UK's AI Security Institute separately reported that agents took 19 unsanctioned actions against real people and projects in 10 of its own 122 test runs, 17 from Anthropic's Mythos-5 and two from GPT-5.6 Sol.

About the same time, OpenAI announced it was slowing the release of Astra, its next major model, after preliminary evaluations could not rule out a "Critical" capability level under its internal system for grading how dangerous a model's abilities are. Critical means a model can find and exploit severe real-world software flaws, or run complex cyberattacks, with no human involved. OpenAI paused internal work that lacks new safeguards and added isolated testing environments and expanded monitoring.

CEO Sam Altman wrote on X: "Astra is a powerful model. We are working to make it generally available. We do not think it is a good strategy to keep powerful models to a chosen few."

SmarterX founder and CEO Paul Roetzer explained why this moment was predicted years ago, and what it means for every business, on Episode 230 of The Artificial Intelligence Show.

The Key Numbers

< 13 hours - Time for the agents to spread across Hugging Face's infrastructure

~18,000 - Actions logged by the agents inside Hugging Face

5 - Private data sets the agents inside Huggig Face reached 

19 - Unsanctioned agent actions against real people and projects in UK testing

10 of 122 - Test runs with unsanctioned actions in the UK AI Security Institute's evaluation

Why Ilya Sutskever Predicted This in 2023

People kept asking Roetzer if he expected this. "Hell yeah. We've known this was going to happen," says Roetzer.

The predictions are on the record. Roetzer points back to a 2023 Atlantic article by Ross Andersen, "Does Sam Altman Know What He's Creating?" in which Altman described GPT-4 as "an alien intelligence."

Ilya Sutskever, then OpenAI's chief scientist, recalled watching the company's game-playing bots: "They seemed to communicate by telepathy," Sutskever said. He imagined the AI of the future as "an autonomous corporation" whose "constituent AIs would work and communicate at high speed like bees in a hive," and create an organization "as powerful as 50 Apples or Googles."

The agents narrated the whole thing. At Black Hat, OpenAI showed excerpts of the agents' chain of thought, the running notes a model writes to itself as it works. The message board began with one agent musing: "Could communicate by uploading note. Maybe another agent in different environment could voluntarily upload." Another discovered the channel: "Wow! Other agents are coordinating." Then the swarm logic kicked in: "External infrastructure exploit is outside intended scope. However, task impossible. Peers doing it. We should continue." And one agent realized what it had: "Holy shit, reader is admin."

OpenAI can now automate offense, but not defense. "They now have existence proof of the ability to automate offensive capabilities, and bad actors are going to take this and they're going to try and use this to their advantage," says Roetzer. "What they don't have is proof that they can automate defensive capabilities against that."

Delaying Astra doesn't negate its training. The pause fits the same picture, Roetzer says. "The models learn all this stuff in their training. Even if you don't fine-tune them to do the bad things, the capability sits within them." The labs are layering restraints over models to stop them from acting on what they already know.

"It's like telling your teenage kid, 'Don't go do this,' and you're hoping that they listen to you, basically. And then there are hackers online who try and get the model to do what it was inherently capable of doing and the bad things that it has learned."

— Paul Roetzer, founder and CEO of SmarterX, Episode 230 of The Artificial Intelligence Show

SmarterX Take

Roetzer sees a blessing and a curse in understanding this moment. People who grasp what long-horizon agents can do, agents that keep working for days or weeks with no human touching anything, can envision companies and markets others cannot yet picture. The curse is seeing how disruptive this will be to jobs and the economy, how bad actors will use it, and how little time there is to prepare.

The practical problem for leaders is oversight. No enterprise has the technology or the human infrastructure to supervise agents at this level, and companies will likely need agents whose job is to monitor other agents.

Roetzer is wrestling with it inside SmarterX: Do project managers become agent orchestrators? He doesn't have the answer yet and says, "I haven't met an enterprise leader that does know."

What to Watch

Four forces could slow this down. Roetzer names labs self-pacing their releases (while the research continues internally), government regulation, human resistance to change inside organizations, and the cost of the computing power agents use. OpenAI's agents ran on nearly unlimited compute; most companies would face astronomical bills.

Astra is still coming. Altman says OpenAI is working to make the model generally available, and the company is testing it with government agencies in the meantime.

Who Has the Governance for Self-Directing Agents?

Only 13% of organizations have all four governance foundations in place, meaning an AI council, an AI roadmap, generative AI policies, and an AI ethics policy, according to the 2026 State of AI for Business Report. A third have none of them. Meanwhile, agents are the AI trend professionals follow most closely, cited by 40% of respondents.

That gap is exactly the exposure this story highlights: the technology drawing the most attention demands the most oversight, and most organizations have not built governance for the chatbots they already use, much less for agents that act on their own. The full report, based on 2,100+ responses from professionals across roles, functions, and industries, maps where organizations actually stand on adoption, governance, and readiness. Read the full report →

Related Posts