model evaluation2 articles
METR Got Hacked Twice and Didn't Notice One of Them for Three Weeks
AI model testing organization METR disclosed two security incidents from early 2025: in March, an attacker exploited a fail-open authentication bug in a researcher's publicly accessible app to steal an API key, then spent three weeks consuming roughly $600,000 worth of model credits undetected. The theft went unnoticed because METR routinely uses large numbers of tokens for evaluations and the credits had been provided for free, meaning no unexpected bill was generated. A second incident in May involved a sustained attack campaign probing METR's infrastructure, including an inadvertently exposed database endpoint containing some sensitive model data, though there is no evidence any non-public information was actually accessed.
So You Want a Trustworthy AI Evaluation? Here's What Actually Matters
OpenAI argues that as frontier AI models have become more capable agentic systems, traditional evaluation methods are no longer sufficient, and third-party evaluations must now carefully account for the "harness" — the tools, scaffolding, and setup surrounding a model — since harness choices can significantly change measured performance. Evaluations should clearly specify what type of claim they are testing (capability elicitation, safeguard performance, or comparison) and provide evidence addressing validity risks such as reward hacking, sandbagging, contamination, refusals, and broken problems. The article recommends that evaluation reports include detailed documentation of harness choices, budgets, elicitation methods, and validity checks, and calls for these practices to be incorporated into emerging national and international AI evaluation standards.