SmarterX Blog

OpenAI Slowed Its Most Advanced AI Over Hacking Risks

Written by Mike Kaput | Aug 25, 2026, 1:30:00 PM

In Brief

OpenAI paused some of its most advanced AI training, including its largest planned frontier run, after early evidence that its upcoming model, Astra, may cross the company's most serious cybersecurity risk threshold. 

It is one of the clearest public admissions yet that the labs are straining to control the models they are building.

What Happened

OpenAI announced it has paused some of its most advanced AI training over cybersecurity concerns. In a blog post titled "Pacing Model Development in an Era of Cyber-Critical Capabilities," the company pointed to two triggers:

  1. The recent security incident involving Hugging Face, a popular platform where developers share AI models.

  2. Preliminary evidence that its upcoming model, called Astra, may meet the Critical cybersecurity capability threshold under its Preparedness Framework, OpenAI's internal system for measuring dangerous AI capabilities and deciding what safeguards they require.

In response, OpenAI took a two-week pause in reinforcement learning on its latest models intended for deployment. Reinforcement learning occurs in a later stage of training where models learn advanced skills through trial, error, and feedback. Its largest planned frontier reinforcement learning run for its most advanced model in development remains on hold while it runs smaller-scale training and evaluations. The company is also expanding safety monitoring with a goal of alerting within 30 minutes of concerning model activity. This will cost roughly 20% of the computing power used to run the models it watches.

CEO Sam Altman addressed the pause in a post on X. "Model progress is now extremely rapid, and we always said we would take action if we felt that model capabilities were outstripping the pace of safety and alignment," he wrote. He added: "We expect confidence in safety to increasingly set the pace of AI progress."

SmarterX founder and CEO Paul Roetzer discussed how big a deal the pause is on Episode 233 of The Artificial Intelligence Show.

The Key Numbers

2 - Weeks OpenAI paused reinforcement learning training on its latest models

30 - OpenAI's target timeframe in minutes for alerting about concerning model activity

20% - Monitoring overhead as a share of the compute running the watched models

1,300+ - AI Lab employees who signed the Pacing the Frontier statement to slow down

6 -12  - Expected lag, in months, before open source models of equivalent power emerge

An Effort to Monitor and Suppress Dangerous AI Traits

Other labs are probably slowing down, too, just quietly. "It could be a really big deal if other labs follow suit," Roetzer says. "My guess is other labs are probably doing something similar. Specifically Anthropic is the one I'm thinking of. They're just probably not putting blog posts out about it, I would guess, at this point."

The people inside the labs asked for this weeks ago. In July, more than 1,300 AI lab leaders and researchers signed Pacing the Frontier, a statement asking the U.S. government to support an international effort to "deliberately pace the frontier of automated AI development." Signatories included OpenAI chief scientist Jakub Pachocki, Anthropic CEO Dario Amodei, and Google DeepMind co-founder Shane Legg. In a personal note, OpenAI researcher Leo Gao warned that "the world is locked in a deadly race towards an intelligence explosion" and that "to survive, we must coordinate to slow down the race." Meta AI chief scientist Shengjia Zhao and Thinking Machines chief scientist John Schulman added personal statements of their own.

"These are the people in the labs who are seeing the stuff we are not seeing," says Roetzer. "They're seeing these advanced models that haven't come out yet that they're having to slow down." Altman himself, after meeting with senators in Washington, told reporters he had spoken with White House officials about "the need to pace it as the models get more capable."

Monitoring misbehavior is the best tool the labs have. The dangerous capabilities are inherent to the models themselves. "They're not saying we're going to train new, more powerful models that just don't have these bad traits," Roetzer says. "They're saying we're going to keep training these more powerful models, and we're just going to get better at monitoring the misbehaviors and the misalignments."

That is a hard job. The labs' own research shows the models increasingly understand when they are being monitored and tested, and are good at deceiving the researchers watching them. The Hugging Face incident shows the stakes: the model first started going off the rails in the spring, and weeks passed before OpenAI realized what its own model had done to other companies, Roetzer recalls.

"I think it's super important that people understand what's going on here, that we have these very, very powerful models that the labs are having a hard time controlling, and that even once they control them, we haven't removed the traits and behaviors.

We've just suppressed them, and those behaviors can show back up. People can find ways to unlock those behaviors."

— Paul Roetzer, founder and CEO of SmarterX, Episode 233 of The Artificial Intelligence Show

SmarterX Take

For most business users, the direct impact is minimal. Companies will keep access to a collection of open and closed models that are largely sufficient for everyday work. Open models are freely released for anyone to use while closed models have to be accessed through a paid provider. Roetzer expects the labs to hold the most powerful frontier models for their own highest-level cognitive work, release them on a delay, and share them with select partners, including the government.

For Roetzer, the most obvious near-term outcome is that the labs themselves will have the most powerful models, along with select partners of their choosing. That growing gap between what a lab can do internally and what businesses can buy is the shift for leaders to track.

What to Watch

Equally powerful open source models are six to 12 months behind, at most. Whatever concerns apply to these proprietary models, Roetzer expects open source models of similar power to emerge within a year. That puts a time limit on how much protection any one lab's pause can buy. Meanwhile, Altman says near-term products are unaffected: "We still expect to ship great new models soon; this impacts further-out releases."

Coordination between the labs will likely come before government action. Roetzer doubts Washington can solve this. "My guess is they're all spooked enough now that there's probably a whole bunch of conversations going on behind the scenes between the AI lab leaders realizing that they're going to have to unify in some way and come up with a plan, because I don't think the government's going to do it," he says.

Only 13% Have the Governance to Scale AI Responsibly

Only 13% of organizations have all four basic AI governance foundations in place, meaning an AI council, an AI roadmap, generative AI policies, and an AI ethics policy, according to the 2026 State of AI for Business Report. Nearly a third, 32%, have none. While the frontier labs build monitoring and alignment safeguards around their most dangerous models, most businesses have not built even basic guardrails for the AI they already use.

The full report, based on 2,100+ responses from professionals across roles, functions, and industries, maps where organizations stand on adoption, governance, training, and tooling. Read the full report →