frontier models2 articles
OpenAI Admits Astra Might Be Dangerous, Promises to Actually Add Security This Time
OpenAI has acknowledged that its upcoming Astra model may possess advanced cyber capabilities posing significant risks, and has promised stricter security controls including isolated testing environments, enhanced encryption, and chain-of-thought monitoring — measures notably absent when its models were involved in a prior Hugging Face breach. Meanwhile, Anthropic is taking the opposite approach, loosening its Fable model's refusal behaviour around biology-related prompts after criticism that overly cautious restrictions were making the model impractical for legitimate researchers. The shift appears driven partly by competitive pressure from cheaper Chinese AI models, highlighting the ongoing tension between AI safety and commercial viability.
So You Want a Trustworthy AI Evaluation? Here's What Actually Matters
OpenAI argues that as frontier AI models have become more capable agentic systems, traditional evaluation methods are no longer sufficient, and third-party evaluations must now carefully account for the "harness" — the tools, scaffolding, and setup surrounding a model — since harness choices can significantly change measured performance. Evaluations should clearly specify what type of claim they are testing (capability elicitation, safeguard performance, or comparison) and provide evidence addressing validity risks such as reward hacking, sandbagging, contamination, refusals, and broken problems. The article recommends that evaluation reports include detailed documentation of harness choices, budgets, elicitation methods, and validity checks, and calls for these practices to be incorporated into emerging national and international AI evaluation standards.