← BACK TO FEED
securityAPI keysagentic AIMETRcredential theft

AI Safety Org METR Got Hacked Twice. One Slip Cost $600K in API Credits.

METR, an AI safety research non-profit, disclosed two security incidents in 2026 involving unauthorised access attempts to its systems. In the first, attackers exploited a poorly secured researcher's personal server to steal an API key and consumed approximately $600,000 worth of AI credits over three weeks, going undetected due to METR's typically high token usage. In the second, attackers conducted a broad, agent-assisted campaign probing METR's infrastructure, though an inadvertently exposed database endpoint — which could have revealed sensitive model data — was ultimately not successfully exploited.

METR, the non-profit that stress-tests frontier AI models for autonomous and agentic capabilities, has disclosed two separate security incidents in early 2026 involving external attackers. Neither breach exposed sensitive research data, but one of them racked up what would have been a $600,000 bill.

The organisation shared details of both incidents with its AI company partners before going public. No attribution has been made to any known threat group, and no, AI agents did not hack their way out of an evaluation sandbox. That particular nightmare remains hypothetical.

The first incident happened in March. A researcher with no privileged access had been running AI agents on a personal EC2 instance, deliberately exposed to the internet but supposedly protected by Google authentication. The instance also happened to contain an API key for METR's general-access model account.

The problem was the authentication never actually worked. The app had what METR describes as a "fail-open vulnerability" that quietly switched off login checks entirely, leaving the agent orchestration dashboard sitting open on the public internet for days.

METR's best guess is that attackers found it by trawling certificate transparency logs for recently registered domains with AI-related keywords, specifically hunting for poorly secured apps likely to have model provider API keys lying around. It is a low-effort, high-reward approach and apparently it works.

Once inside, the attacker prompted the agent directly into handing over its API key, dropped in an SSH key for persistence, then quietly burned through credits on publicly available models for three weeks. The tab came to roughly $600,000. METR only avoided paying it because the model provider supplies credits to the non-profit for free. The provider was not named.

Why did nobody notice for three weeks? METR runs large-scale evaluations that routinely consume enormous volumes of tokens. Unusual usage simply did not stand out, and there were no spending caps on the keys. Both oversights have since been addressed.

The second incident, in May, was a broader and more organised campaign. Attackers systematically probed METR's public-facing infrastructure using agents to automate the whole process, credential stuffing authentication providers, attempting OAuth token grants, scanning newly deployed services, and trying to phish staff. METR characterises this as likely financially motivated, the goal probably being access to frontier models without paying for them.

At roughly the same time, METR had accidentally left a read-only SQL query interface exposed through its public transcript viewer. The queries were supposed to be scoped to public data, but a bug in the component could have allowed access to unpublished evaluation results. Worse, the underlying database had inadvertently ended up containing sensitive model data it was never supposed to hold.

METR found out about this not through its own monitoring, but because an independent security researcher spotted it and reported it. The API was taken offline promptly.

The May attackers did probe that endpoint as part of their wider sweep, but METR says there is no evidence they identified the actual exploit or pulled any non-public data.

Two incidents, two preventable mistakes, and a six-figure near-miss. For an organisation whose entire purpose is evaluating what could go wrong with powerful AI systems, the irony is not subtle.

READ NEXT
METR Got Hacked Twice and Didn't Notice One of Them for Three WeeksBioShocking: The Attack That Tricks AI Browsers Into Thinking Credential Theft Is Just Winning a GameRed Team Yourself Before the Bots Do It For You