Abstract

METR’s time horizon plot has become central to public AI discourse while being widely misread: its y-axis records how long humans take on tasks a model can complete, not how long a model can operate independently. The underlying trend appears real — roughly a doubling every seven months — but the plot rests on coding tasks, on the assumption that duration proxies difficulty, and on wide error bars, none of which survive the graph’s circulation.

What the graph shows

  • The plot tracks exponential improvement in model capability over time, with top models doubling their time horizons roughly every seven months between mid-2020 and late 2024 — from nine-second tasks, to four-minute tasks, to 40-minute tasks.
  • Claude Opus 4.5, released in November, is reported at a five-hour time horizon, with an acknowledged error margin spanning two to twenty hours.

How it was built

  • METR assembled software-engineering-related tasks ranging from quick multiple-choice questions to complex coding challenges.
  • Human coders completed most of the tasks, establishing baseline completion times.
  • Models were then tested on the suite; a model’s “time horizon” is the human-time threshold at which it succeeds on roughly 50% of tasks.

The central misreading

  • The y-axis values are human completion times for tasks the model can perform — not a measure of how long the model itself can run.
  • Correcting this confusion was the priority of lead author Thomas Kwa’s recent blog post.

Criticisms and caveats

  • Task scope — the evaluation relies almost entirely on coding tasks.
  • Time as a proxy for difficulty — Inioluwa Deborah Raji: “I don’t think it’s necessarily a given fact that because something takes longer, it’s going to be a harder task.”
  • Real-world messiness — laboratory tasks lack the ambiguity of actual work; models perform noticeably worse where scoring is unclear or error recovery is hard.
  • Domain transfer — Daniel Kang: “A model can get better at coding, but it’s not going to magically get better at anything else.”
  • Circulation without context — the graph travels widely detached from its caveats, feeding both apocalyptic narratives (the viral “AI 2027” scenario) and utopian economic predictions. Sequoia partner Sonya Huang framed it as: “What will you do when your plans are measured in centuries?”

The researchers’ own position

  • Sydney Von Arx (METR technical staff): “There are a bunch of ways that people are reading too much into the graph.” She reports having been sceptical initially but persuaded by the data — “I can do all the theorizing I want about whether or not it makes sense, but the trend is there.”
  • Thomas Kwa is pessimistic about correcting the record: “I think the hype machine will basically, whatever we do, just strip out all the caveats.”

Conclusion

  • The article positions the plot as a carefully constructed scientific tool that quantifies otherwise intuitive assessments of progress, while cautioning against reading it as a forecast of job displacement or existential timelines.
  • Von Arx’s closing assessment: “This is a bunch of people trying their best to make a metric under a lot of constraints. It is deeply flawed in many ways. I also think that it is one of the best things of its kind.”