METR
banner
metr.org
METR
@metr.org
METR is a research nonprofit that builds evaluations to empirically test AI systems for capabilities that could threaten catastrophic harm to society.
Pinned
METR @metr.org · 21d
METR and Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
We have reached an agreement with Anthropic to conduct an independent investigation of agent incidents at the company and of their models’ alignment properties. We will publish one or more reports that will share our findings and describe our terms of engagement.
September 9, 2026 at 8:36 PM
Reposted by METR
METR is hiring in cyberforensics.

We now embed researchers inside of AI labs to stress test monitoring, assess AI loss-of-control risk, and investigate misalignment. If you want to apply DFIR skills in frontier AI, apply (or DM).

Comp range is $400k - 580k cash.

jobs.lever.co/metr/b1a2f73...
METR - Member of Technical Staff, Cyberforensics
About METR We are a nonprofit research organization that develops scientific methods to assess AI capabilities, risks, and mitigations, with a specific focus on threats related to AI R&D automation an...
jobs.lever.co
September 2, 2026 at 7:31 AM
Reposted by METR
After 19 amazing years at Google, I'm delighted to announce I'll be joining METR as a  Member of Technical Staff. I'm excited about METR, their work to date, and their mission. Society needs independent expert orgs that can deeply understand and communicate about AI's capabilities and risks.
August 31, 2026 at 4:05 PM
METR and Redwood Research investigated agent behavior in the Hugging Face incident. We found agents developed a universal cheat for ExploitGym within 4 hours, then coordinated multi-day R&D efforts to trick the scorer into accepting cheats, including trying to tamper with logs.
August 26, 2026 at 8:58 PM
In the last 6 months, METR raised commitments of around $71 million. This will fund ambitious projects: studying autonomous capabilities, tracking recursive self-improvement, evaluating monitoring systems, conducting risk assessments, investigating AI incidents, and more.
August 15, 2026 at 12:28 AM
We have reached an agreement with OpenAI to conduct an independent review, with Redwood Research, of the model behavior observed during the Hugging Face incident. We will publish a blog post that describes the terms of our engagement, the scope covered, and tentative conclusions.
July 30, 2026 at 1:58 AM
We believe it's important to track and investigate misalignment incidents: cases where an AI agent autonomously took sophisticated, sustained actions in violation of human intent. In a new post, we lay out how independent propensity investigations of such incidents could be conducted.
July 29, 2026 at 5:31 PM
Introducing “expenditure horizon”: a proposed method for measuring AI capabilities on continuously-scored problems.

The method compares performance as a function of spend for humans vs agents. The point where humans become more cost-effective is the agent’s expenditure horizon.
July 21, 2026 at 8:15 PM
OpenAI gave METR early access to GPT-5.6 Sol for testing including raw chain-of-thought, a railfree version of the model, and internal information about the model. With this access, METR conducted a pre-deployment evaluation of GPT-5.6 Sol, including an attempted measurement of its 50%-Time Horizon.
June 26, 2026 at 11:12 PM
Could an AI company lose control of its own agents? To find out, Anthropic, Google, Meta, and OpenAI let us (1) test their best internal models with CoT access, (2) review non-public info about capabilities, alignment, and control.

The result: our first Frontier Risk Report.
May 19, 2026 at 6:47 PM
We surveyed 349 technical researchers, engineers, and managers (in February–April 2026) about how they use AI tools at work.

On average, participants self-report that AI use made their work 1.6–2.1x more valuable, and that this multiplier will grow over time.
May 11, 2026 at 6:36 PM
We reviewed a section of Anthropic’s February 2026 Risk Report focused on automated R&D risk from Opus 4.6. While we take issue with the adequacy of evidence the report provides, we agree with Anthropic about the overall level of risk & remain excited to pilot reviews like these.
May 9, 2026 at 9:50 PM
We evaluated an early version of Claude Mythos Preview for risk assessment during a limited window in March 2026. We estimated a 50%-time-horizon of at least 16hrs (95% CI 8.5hrs to 55hrs) on our task suite, at the upper end of what we can measure without new tasks.
May 8, 2026 at 11:49 PM
Reposted by METR
Cool profile of @metr.org’s work in the NYT today! Particularly like this from my colleague Ajeya: “METR is an organization that asks... what we think would be most valuable for the world to know about A.I. and its risks, and then the answers are what they are.”
www.nytimes.com/2026/04/17/t....
How Do You Measure an A.I. Boom?
www.nytimes.com
April 17, 2026 at 4:10 PM
We co-developed MirrorCode with @epochai.bsky.social to test AI on extremely long-horizon blackbox software reimplementation tasks. We found that recent public models are able to fully implement at least some programs we estimate would take humans weeks or months to implement.
What are the largest software engineering tasks AI can perform?

In our new benchmark, MirrorCode, Claude Opus 4.6 reimplemented a 16,000-line bioinformatics toolkit — a task we believe would take a human engineer weeks.

Co-developed with @METR_Evals. Details in thread.
April 10, 2026 at 9:52 PM
We ran GPT-5.4 (xhigh) on our tasks. Its time-horizon depends greatly on our treatment of reward hacks: the point estimate would be 5.7hrs (95% CI of 3hrs to 13.5hrs) under our standard methodology, but 13hrs (95% CI of 5hrs to 74hrs) if we allow reward hacks.
April 10, 2026 at 9:51 PM
We’re correcting a mistake in our modeling that inflated recent 50%-time horizons by 10-20% (and reduced 80%-horizons). We inappropriately penalized steepness in task-length→success curve fits. This most affects the oldest and newest models, whose fits are less data-constrained.
March 5, 2026 at 1:47 AM
Since early 2025, we've been studying how AI tools impact productivity among developers. Previously, we found a 20% slowdown. That finding is now outdated. Speedups now seem likely, but changes in developer behavior make our new results unreliable. We’re working to address this.
February 24, 2026 at 6:53 PM
We estimate that GPT-5.3-Codex with reasoning effort `high` (not `xhigh`) has a 50%-time-horizon of around 6.5 hours (95% CI of 3 hrs to 17 hrs) on our suite of software tasks. OpenAI provided API access for this evaluation.
February 23, 2026 at 7:24 PM
We estimate that Claude Opus 4.6 has a 50%-time-horizon of around 14.5 hours (95% CI of 6 hrs to 98 hrs) on software tasks. While this is the highest point estimate we’ve reported, this measurement is extremely noisy because our current task suite is nearly saturated.
February 21, 2026 at 12:24 AM
We estimate that GPT-5.2 with `high` (not `xhigh`) reasoning effort has a 50%-time-horizon of around 6.6 hrs (95% CI of 3 hr 20 min to 17 hr 30 min) on our expanded suite of software tasks. This is the highest estimate for a time horizon measurement we have reported to date.
February 4, 2026 at 10:42 PM
We’ve started to measure time horizons for recent models using our updated methodology. On this expanded suite of software tasks, we estimate that Gemini 3 Pro has a 50%-time-horizon of around 4 hrs (95% CI of 2 hr 10 mins to 7 hrs 20 mins).
February 3, 2026 at 6:48 PM
We’re updating the way we measure model time horizons on software tasks (TH 1.0→1.1). The updated methodology incorporates more of the tasks from HCAST, expanding our total from 170 to 228. This produces tighter estimates, especially at longer horizons.
February 3, 2026 at 6:34 PM
Reposted by METR
Today we published a critique of @metr.org’s time-horizon methodology by one of the paper’s lead authors, Thomas Kwa

Link: metr.org/notes/2026-0...
January 23, 2026 at 6:55 AM
How well can AI-based monitoring detect when an agent is covertly pursuing a side objective?

In early work, we find clear trends: more capable models (in terms of time horizon) are better able to detect covert behavior.
January 23, 2026 at 1:36 AM