The useful question is changing.

Instead of asking only “How smart is AI?”, ask: How much of a workflow can AI complete reliably before a human has to take over?

What the number actually means

METR estimates a model's task-completion time horizon by looking at tasks that human experts take different amounts of time to complete. A 50% horizon of 12 hours means the fitted model predicts roughly a 50% chance of succeeding on tasks in the evaluation suite that take a skilled human about 12 hours.

That is not the same as an AI running continuously for 12 hours. It is also not evidence that the system can replace a worker for a 12-hour shift. METR's suite is dominated by relatively well-specified software, machine-learning and cybersecurity work, and its researchers report weaker performance on messier tasks.

KNOWN: autonomous task capability is getting longer

METR's February–March 2026 frontier-risk assessment placed the public frontier at roughly a 12-hour 50%-success horizon and roughly a 1.5-hour 80%-success horizon. Its live time-horizon project describes a longer-run pattern of rapidly increasing autonomous capability across frontier models.

The important change is not that AI suddenly became an autonomous employee. It is that agents can complete larger chunks of structured technical work before a person has to intervene.

UNCERTAIN: benchmark capability is not workplace deployment

Real work includes missing context, unclear goals, politics, customer preferences, security constraints, changing requirements and consequences for mistakes. Those conditions are difficult to reproduce in a benchmark.

Reliability also matters more as tasks get longer. A system that succeeds half the time may be useful for experiments or supervised workflows while being unacceptable for an unattended production process.

Three plausible work scenarios

SCENARIO 01

AI stays an assistant

Agents handle bounded tasks, but humans continue to break down work, resolve ambiguity and integrate the result.

SCENARIO 02

Workflow compression

One person supervises several agents and handles work that previously moved through multiple people or handoffs.

SCENARIO 03

Project-level autonomy

Agents eventually execute substantial projects with humans intervening mainly at decision, review and exception points.

Our read

Workflow compression deserves more near-term attention than the idea of every profession suddenly disappearing. If agents can take on longer sequences of work, companies may reorganize teams before they eliminate entire occupations. The first visible effect could be fewer handoffs, higher output per worker and changing expectations for junior roles.

This is an interpretation, not a measured forecast. It should change as new labor-market and deployment evidence arrives.

Signals to watch

01Does the 80%-reliability horizon rise along with the headline 50% horizon?
02Do agents improve on messy, ambiguous tasks rather than only clean benchmarks?
03Do companies report smaller teams producing the same output?
04Does entry-level hiring weaken specifically where agentic workflows are deployed?

How to prepare without predicting an AGI date

Understand the workflow

Learn what happens before and after your current task so your value is not tied to one narrow step.

Practice delegation

Use AI for bounded pieces of real work, then learn how to provide context, constraints and acceptance criteria.

Get good at verification

As output gets cheaper, the ability to recognize a wrong, unsafe or low-quality result becomes more valuable.

Primary sources

Research note: This article separates measured benchmark evidence from our interpretation and future scenarios. We update claims when stronger evidence becomes available.
← Back to The Neural Tide