Abstract

An updated version of METR’s time horizon estimates, built on a larger task suite and open-source evaluation infrastructure, produces estimates that generally fall within the previous confidence intervals — but yields a faster measured trend for recent models, with the doubling time for post-2023 models falling from 165 to 131 days.

What changed

Task suite

  • Expanded from 170 to 228 tasks: 73 new tasks added from HCAST, 15 removed, 53 updated.
  • Long-duration tasks (8 hours and above) doubled from 14 to 31, which tightens confidence intervals for the most capable models.

Infrastructure

  • Evaluation infrastructure migrated from METR’s in-house Vivaria to the open-source Inspect framework.
  • The migration produced only marginal differences in results while improving standardisation.
  • Full window (2019–2025) — doubling time remains approximately 196 days (about 7 months).
  • Post-2023 models — doubling time falls to 131 days under TH1.1, against 165 days under TH1.
  • Post-2024 models — further acceleration to 89 days.

Per-model changes

  • Claude Opus 4.5 rose 11%.
  • GPT-5 rose 55%.
  • Older GPT-4 versions fell 35–57%.

Methodological caveat

  • The authors note the sensitivity of results to task composition: the updated suite samples from a slightly different distribution of difficulty.
  • This is described as a natural consequence of improving the suite in the absence of rigid selection criteria.