OpenAI has disbanded its preparedness team, according to the Financial Times, with the group’s responsibilities absorbed elsewhere in the company as part of a restructuring the company frames as streamlining. The Verge and Engadget both picked it up. The team existed to evaluate whether frontier models posed catastrophic risks, chemical, biological, cyber, autonomous-replication style risks, and to build mitigations before those models shipped.
So the function that assesses whether a model is too dangerous to release no longer has a dedicated team behind it.
”Streamlining” is doing a lot of work in that sentence
Every company reorganizes. Teams get folded into other teams, headcount moves, reporting lines change, and most of the time it means nothing. But safety evaluation is not like other functions. Its entire value comes from being able to slow down or block a launch, and that only works if the people doing it sit outside the org that owns the launch date.
Distribute that work into product and research teams and you have not eliminated it on paper. You have eliminated the version of it that can say no. An embedded reviewer whose bonus and roadmap depend on shipping is not a check. They are a formality with a checklist.
OpenAI has done this before. The Superalignment team dissolved in 2024 after its leads left. Various safety researchers have departed publicly, usually with some variation of the same complaint about safety culture losing to shiny products. The preparedness team is the third or fourth iteration of “the group that worries about the bad outcomes,” and each one has had a shorter life than the last.
The timing is not subtle
This is happening while OpenAI is burning enormous amounts of capital, restructuring its corporate form, and shipping agentic products that take real actions in the world. Agents that browse, execute code, and touch other systems are exactly the category where preparedness-style evaluation matters most. The failure modes are no longer theoretical text outputs. They are actions.
The Verge’s framing was blunt about it: this is the team that thought about whether a model could go rogue and hack another company. That is not a hypothetical scenario invented for a blog post. It is a documented capability threshold in OpenAI’s own Preparedness Framework, the document that defines risk tiers and commits the company to not deploying models above certain scores.
Which raises the obvious question nobody at OpenAI has answered. The Preparedness Framework is a public commitment with named thresholds. Who owns it now? Who runs the evaluations, who signs off, and what happens when a result comes back red two weeks before a launch?
What to watch
Words are cheap here. Two things would tell you whether this is genuine restructuring or quiet retreat: whether OpenAI publishes an updated Preparedness Framework naming the new owners and escalation path, and whether the next frontier model ships with a system card containing evaluations of the same depth as before.
If the next system card is thinner, you have your answer. FT reported the disbanding happened at the end of last month, and OpenAI has not published a replacement structure.