← BackMETR (Model Evaluation & Threat Research)
research organizationCredibility: 82%
Why this score? Independent nonprofit third-party evaluator of frontier AI dangerous capabilities; formerly ARC Evals. Author of the task-horizon study; partners with labs on pre-deployment evals. High independent-evaluator credibility.
Tracked Statements (1)
Context: METR’s own extrapolation from its measured six-year doubling trend; not yet resolvable. METR notes a 10x error in the absolute measurement would shift arrival estimates by only ~2 years, and that the SWE-Bench-Verified subset doubled even faster (under 3 months).