Zero Day Room
Live
Threats

METR API Key Theft Leads to $600,000 AI Credit Burn

Attackers stole an API key from AI safety research group METR and used it for three weeks, consuming model credits valued at approximately $600,000.

Attackers stole an API key from AI safety research group METR and used it for three weeks, consuming model credits valued...

Attackers stole an API key from the AI safety research organization METR and used it for three weeks to consume model credits with an estimated commercial value of $600,000. METR disclosed this incident from March, alongside a separate attack in May, in a security update published on August 31, stating it found no evidence that sensitive information was accessed.

The credits had been supplied free of charge by an unnamed model developer, meaning the $600,000 figure represents their market value rather than a direct financial loss to METR. The organization emphasized that the disclosure concerned external human attackers, not AI agents acting within its own evaluations, where an initial scan found no signs of agents hacking third parties.

Vibe-Coded App Exposed API Key

The breach began in March when a METR researcher ran agents on a personal Amazon EC2 instance. This instance was made publicly accessible behind Google authentication via a "vibe-coded" application that held an API key for METR's public models account.

A fail-open flaw silently disabled the authentication, leaving the system exposed for several days. METR suspects the attacker found the instance by mining certificate transparency logs for recently registered sites containing high-signal terms related to language models and AI agents.

The attacker prompted an agent to reveal the model provider's API key, added an SSH key for persistence, and then used the stolen credentials. Over the following three weeks, they consumed large volumes of model credits. METR said the illicit usage was difficult to distinguish from legitimate evaluation activity, as its researchers routinely generate high volumes of model traffic. The organization also had no mechanism to cap spending on keys associated with free credits.

In response, METR revoked the researcher's access, rotated all credentials, wiped the affected laptop, and alerted the model developer. It later added spending alerts to keys where such controls were possible.

Second Attack Probed Public Infrastructure

In early May, METR was tipped off that it was being targeted by financially motivated attackers who may have been seeking access to frontier AI models. These attackers heavily used automated agents to discover vulnerabilities.

Their techniques included credential stuffing, attempts to grant OAuth tokens, scanning of new services, and attempts to phish staff. METR also disclosed an unrelated security flaw in its public transcript viewer, which inadvertently exposed a read-only SQL query mechanism.

A bug in this interface could have been exploited to access unpublished evaluation data. Furthermore, the database had accidentally been loaded with sensitive model data it was not designed to hold. An independent researcher disclosed this flaw, leading METR to take the interface offline and pay a bug bounty. The organization stated the May attackers probed this endpoint but did not appear to discover or exploit the specific bug.

METR's Security Improvements

METR described its wider security posture as accurate to July 30. As a result of these incidents, the organization now runs public-facing applications in an environment that is architecturally separated from its internal infrastructure. This measure aims to prevent future cross-contamination between exposed services and core systems.

The update from METR serves as a case study in the challenges of securing AI research environments, particularly when handling high-value model credits and managing the blend of automated agent use and human error. The lack of spending caps on free-credit keys was a specific weakness highlighted by the March theft.

Related coverage

More from Threats