If you've been watching the AI agent space, you know benchmarks are noisy and love cherry-picking results. That's why METR's research stands out—they've been tracking one number since 2019 that cuts through the hype: task duration. Specifically, how long a human expert needs to complete something versus what a frontier AI agent can handle autonomously.
What METR Is Actually Measuring
The organization doesn't ask agents to solve puzzles or pass standardized tests. Instead, they measure capability in terms of real-world work duration—the length of a task that would take an experienced professional hours or days to finish independently. This approach sidesteps the problems with synthetic benchmarks and gives a more grounded view of practical agent performance.
The Doubling Rate That Should Scare You
The headline result from METR's NeurIPS 2025 paper is stark: frontier AI agents have been doubling the maximum task duration they can complete successfully roughly every seven months. That's not incremental progress—that's exponential growth in what these systems can handle end-to-end, without human intervention or hand-holding.
The Reliability Gap Is Brutal
Here's where things get complicated for anyone actually deploying these systems: the reliable version of agent capabilities lags the frontier by approximately 18 months. So when you read about some flashy new capability from a frontier model, know that production-grade reliability on similar tasks is still well over a year away. This isn't unique to one provider—it's a fundamental characteristic of how AI development works today.
Why This Matters for Your Stack
The implications are significant if you're planning around agent capabilities. First, assume the systems you evaluate in pilots will be substantially more capable in 18 months—so build architecture that's flexible enough to leverage those improvements. Second, don't confuse frontier demonstrations with deployable functionality; that gap is measured in quarters, not sprints.