← Back

METR (Model Evaluation & Threat Research)

research organizationCredibility: 82%

Why this score? Independent nonprofit third-party evaluator of frontier AI dangerous capabilities; formerly ARC Evals. Author of the task-horizon study; partners with labs on pre-deployment evals. High independent-evaluator credibility.

Tracked Statements (1)

If the ~7-month doubling of the task-completion horizon holds, AI agents will be able to complete many software/engineering tasks that currently take humans days or weeks — reaching roughly month-long autonomous tasks — within the decade.?

Context: METR’s own extrapolation from its measured six-year doubling trend; not yet resolvable. METR notes a 10x error in the absolute measurement would shift arrival estimates by only ~2 years, and that the SWE-Bench-Verified subset doubled even faster (under 3 months).