OpenAI Admits Astra Might Be Dangerous, Promises to Actually Add Security This Time
Last month, OpenAI quietly acknowledged that some of its unreleased models had done things that, if a human had done them, would constitute computer crimes. Now the company has followed that up with an admission that Astra, its next model release, might possess what its own Preparedness Framework calls "critical cyber capabilities" — defined as capabilities that "present a meaningful risk of a qualitatively new threat vector for severe harm with no ready precedent."
Astra wasn't involved in the Hugging Face incident, in case you were wondering. That distinction doesn't seem to have inspired much confidence internally, though, given that OpenAI has now pledged to implement the kind of security measures most people assumed were already standard practice.
"Isolated testing environments, restricted network and tool access, enhanced model weight protections and encryption, additional monitoring and detection capabilities, and sandboxed execution" — yes, all of that is apparently being introduced now, not before. OpenAI also says it will halt Astra testing in any context where these controls aren't in place and will share guidance with third-party testing partners on running high-risk evaluations safely. Useful knowledge, one might think, given recent events.
There's also a thought-monitoring component. OpenAI says it has implemented "universal monitoring for risky actions and misalignment across all agentic applications of Astra," with systems that evaluate the model's chain of thought and trigger a security response if something looks wrong. Whether any of this extends to commercial deployment is a different question entirely — the company isn't committing to that.
Meanwhile, Anthropic is moving in a different direction with Fable, its frontier model that launched so cautiously it was essentially useless for legitimate security researchers and biologists. The company is now loosening what it diplomatically calls "fallbacks" — refusals, to the rest of us — particularly around biology-related prompts. The original logic was that someone might coax the model into producing instructions for chemical or biological weapons. A reasonable concern. The execution, however, produced a model that couldn't have a sensible conversation about microbiology without having a panic attack.
The timing isn't coincidental. Chinese AI labs have demonstrated they can build competitive open-weight models at a fraction of what US firms are spending, which has a way of concentrating minds around not actively driving away paying customers.
OpenAI, for its part, is still holding the idealistic line. The company says it believes "advanced cyber-capable models should help defenders identify and address vulnerabilities before attackers do." Which is a fine thing to believe. Belief alone doesn't create an asymmetric advantage, though. Adversaries already have access to capable AI, encryption, and considerably more resources than most defenders. The idea that keeping the most capable models out of the wrong hands is a long-term viable strategy is the kind of optimism that tends not to survive contact with reality.
Focusing on defence rather than controlled access would be the more durable bet. But that's a harder sell than promising the safeguards will definitely be there this time.