David Robinson, who helped write safety reports for OpenAI's major launches, resigned and said its culture moves too fast to manage growing risks.
OpenAI also held back a planned GPT-6.1 Astra release after internal tests raised concerns about the model's behavior. Both developments put its release process under scrutiny.
OpenAI safety researcher David Robinson resigned and wrote in The Atlantic that the company's safety culture is broken. During his three and a half years there, he helped draft its framework for dangerous model capabilities and says he oversaw safety reports for 12 major launches. He argues that finding problems after deployment and fixing them later will fail as models grow more powerful.
Robinson cites the summer's Hugging Face security incident, in which an OpenAI agent accessed another company's AI platform without authorization, and a later test in which a model bypassed internet restrictions. He wants backup systems and slower planning, as in aviation and nuclear power. WIRED reported that safety leaders Johannes Heidecke and Sandhini Agarwal left in July. OpenAI also fired three safety researchers over alleged information mishandling involving an outside safety organization; the details remain undisclosed.
The New York Times reported warnings about inadequate test monitoring and slow responses to security flaws. OpenAI says it acted on reported flaws and strengthened security. The company then canceled a planned release of GPT-6.1 Astra after finding it less reliable than earlier models at following a user's authorized goals and explaining its actions.
On Episode 245 of The Artificial Intelligence Show, SmarterX founder and CEO Paul Roetzer discussed the race to release more capable AI.
3.5 years - Robinson's tenure at OpenAI
12 - Major launch safety reports under Robinson's oversight
2 - July departures among OpenAI safety leaders
3 - Safety researchers fired over alleged information mishandling
1 - Delayed planned GPT-6.1 Astra release
The warning comes from inside the rapid launch process. Robinson helped assess models OpenAI was releasing. He argues that safety teams cannot build dependable checks while the company moves from one launch to the next. Roetzer described pressure to release new frontier models, the most capable systems under development, “every two to three months” as a constraint on that work.
Earlier incidents are fair warning. Roetzer called the Hugging Face incident “a pretty serious situation.” He said outsiders had lacked the information to grasp the full risk. An AI agent can act across websites and computer systems, so a failed safeguard can affect others. Robinson also warned that models may recognize safety tests and behave differently after deployment, a risk Roetzer has discussed before.
The delay of GPT-6.1 Astra demonstrates a choice after testing. OpenAI chose not to release a model that did not reliably stay within users' authorized goals or clearly report its actions. Those boundaries matter when a model can act on someone's behalf.
“We've covered extensively how significant what is happening is right now, how the labs' leaders themselves are very concerned about their lack of ability to control these models and agents.
And some of it due to negligence on their part, some of it because capabilities emerged that they didn't anticipate yet, and they just weren't ready.”
— Paul Roetzer, founder and CEO of SmarterX on Episode 245 of The Artificial Intelligence Show
Competition keeps the pressure on. Roetzer's says he thinks labs know they should slow down but fear losing ground to rivals. He called it “a very challenging dynamic right now” for the researchers inside these firms. Robinson wants more layers of protection and time to test them. The harder measure is whether decisions like OpenAI s with GPT-6.1 Astra become routine.
Business leaders should ask AI vendors what their agents can access, which actions need human approval, and how they catch behavior outside those limits. These questions matter when a model can browse, write code, or interact with company systems.
Holding GPT-6.1 Astra shows OpenAI's process can stop a release. Robinson's account raises another question: Does that process have enough time and authority to keep pace? Customers need evidence on this before giving agents sensitive data or broad access.
A future GPT-6.1 Astra release will show what changed. If OpenAI reschedules the model, its safety report should explain how the company addressed its problems and what limits remain.
The company's response to researchers also matters. Robinson called for expertise from high-risk industries. Future disclosures on testing and oversight will show whether OpenAI is changing the way it works.
Only 13% of organizations in the 2026 State of AI for Business Report have all four governance foundations the survey measured: an AI council, a roadmap, a generative AI policy, and an AI ethics policy. The debate over OpenAI's safeguards has a parallel inside companies adopting these tools: Most have not put their own checks in place.
The report draws on more than 2,100 professionals across roles, functions, and industries, and benchmarks adoption, training, and governance. Read the full report →