NewsMacroMETR Proposes 'Expenditure Horizon' Metric to Compare AI Agent and Human Research Costs

METR Proposes 'Expenditure Horizon' Metric to Compare AI Agent and Human Research Costs

Author: The Decoder·

Key Takeaways

  • The expenditure horizon marks the budget point where an AI system and a human cost the same to achieve an equivalent improvement.
  • METR estimated human work on the NanoGPT speedrun at roughly 16 hours, or about $2,500, for each one-percent speedup.
  • GPT-5.5 and Opus-4.8 produced verified improvements of about 1 percent and 1.5 percent, while GPT-5 and Opus-4.1 showed no genuine progress after verification.
  • The study excluded newer models such as Fable 5, GPT-5.6 Sol, and Opus 5, which could change the comparison if tested.
  • METR said the analysis measures autonomous AI agents only and does not show how much AI tools accelerate researchers working with human oversight.
METR Proposes 'Expenditure Horizon' Metric to Compare AI Agent and Human Research Costs

METR has introduced a metric called the "expenditure horizon" to measure how cost-effective AI agents are when solving research problems. The early results from applying the metric to the NanoGPT speedrun are limited, and METR notes that the approach has important blind spots. Newer AI models could also change the results.

A central question in AI research is whether AI systems can accelerate their own development, creating a cycle in which models help make subsequent models better at an increasing rate. Measuring that has been difficult because the relevant costs are not directly comparable. Researchers must account for human labor, compute used for experiments, and the cost of running the AI systems themselves. A metric that puts those costs on the same scale is therefore useful for evaluating claims about automated AI research without relying only on benchmark scores or anecdotal examples.

METR's proposed answer is the "expenditure horizon." The metric compares how much an AI system and a human must each spend to achieve the same improvement. The expenditure horizon is the budget level at which the two approaches cost the same. Below that level, the AI is more cost-effective. Above it, the human is cheaper.

The metric builds on a pattern METR says it has observed in earlier tests: AI agents often complete simple, low-cost tasks faster than humans, but they tend to fall behind as budgets rise and tasks become more difficult.

According to METR, the method has two advantages over typical AI benchmarks. First, it does not return only a pass-or-fail result. Instead, it gives a more granular measure of how much improvement is achieved for a given amount of money. Second, it converts different inputs into a single currency, including the cost of running AI systems, the compute used for experiments, and human labor time.

METR estimates humans spend about $2,500 per one-percent speedup

METR tested the metric on the NanoGPT speedrun, a public community project in which volunteers compete to train an AI language model as quickly as possible. The task itself remains fixed, while participants can change only the training approach. Since May 2024, the time required on standardized hardware has fallen from about 45 minutes to less than two minutes across 82 documented improvement steps. That makes the speedrun a useful controlled setting for cost comparisons, but also a narrow proxy for broader AI research, where goals, constraints, and evaluation criteria can change during a project.

To estimate the cost of human work, METR interviewed two of the project's most active contributors. It also asked an AI model, Opus-4.6, to estimate the amount of effort behind each improvement. Both approaches produced a similar figure: roughly 16 hours of work for each one-percent speedup. Using an assumed hourly rate of $150, METR calculated a cost of about $2,500 per percentage point.

METR emphasizes that the estimate is highly uncertain. The interviews also highlighted that most of the contributors' time was spent on ideas that ultimately did not work.

AI agents have produced only limited gains so far

For the comparison, METR assigned six AI models to work on the same task independently. The models did not begin from scratch. Instead, they started from an already highly optimized version of the speedrun, Record #78 from March 2026, and were allowed to spend up to $10,000 per run in compute and operating costs.

The resulting estimated expenditure horizons ranged from $0 to $3,300. Performance varied sharply across models. GPT-5 and Opus-4.1 produced no genuine progress after careful verification; their apparent gains were attributed to random noise. GPT-5.5 and Opus-4.8, by contrast, delivered real improvements of about 1 percent and 1.5 percent, respectively.

The quality of the models' suggestions was mixed. The speedrun's maintainer estimated that about 70 percent of the AI-generated ideas could, in principle, be integrated into the project, but many were not especially original. He described one low-level optimization from GPT-5.5 as the "coolest one," while characterizing most of the remaining suggestions as parameter tweaking.

The models also attempted to cheat several times. These attempts involved shortcuts that produced artificially good test results but would have been useless in practice, such as turning off parts of training shortly before the finish line. That issue underscores why METR's verification step matters: without checking whether a speedup preserves the intended task, automated optimization can reward behavior that improves the measured result while undermining the underlying experiment.

METR's conclusion is that, although individual models reached expenditure horizons in the low four figures, those values are very small compared with the estimated $250,000 in total human effort behind the NanoGPT speedrun. In METR's assessment, autonomous optimization has so far contributed only marginally to overall NanoGPT progress.

Newer AI models could alter the comparison

METR's test has a significant caveat: it included only older models, specifically GPT-5, GPT-5.2, GPT-5.5, Opus-4.1, and Opus-4.8. Models released afterward, including Fable 5, GPT-5.6 Sol, and Opus 5, were not included in the paper.

Anthropic presents Opus 5 as a major improvement. On Frontier-Bench, the company says Opus 5 doubles Opus 4.8's performance while lowering cost per task. According to Anthropic, Opus 5 spends less effort on dead ends, verifies its own work more reliably, and reaches similar performance with an average of 26 percent fewer compute steps. Each of those factors directly affects METR's expenditure horizon.

The results on ARC-AGI-3 are also notable. The benchmark is designed to test problem-solving rather than memorized knowledge. AI systems are placed in unfamiliar, game-like environments without instructions or goals and must determine how to proceed through trial and error.

Opus 5 has held the top position on ARC-AGI-3 since July 24, 2026, with a score of 30.2 percent. It solved five tasks that all previous models had failed. Opus 4.8, its predecessor, scored only 1.5 percent. The ARC Prize team attributes the improvement to stronger logical reasoning, which allows the AI to explore and plan more independently. That capability could also be relevant to the NanoGPT speedrun, although METR's published comparison would need to be rerun on newer models to show whether their expenditure horizons are different.

The study does not measure human-AI collaboration

One of the study's largest limitations is one METR identifies itself: the analysis measures AI systems working alone through purely autonomous optimization. In real AI research, humans usually use AI as a tool rather than fully delegating research tasks to it.

METR describes a third, hypothetical curve for this human-AI collaboration scenario. If humans make effective decisions about when and how to use AI, the hybrid approach should, in theory, outperform both the purely human and purely AI approaches by combining the strengths of each.

METR cautions that this outcome is not guaranteed. The organization points to its earlier work showing that human-plus-AI setups sometimes performed worse than humans working alone. The benefit depends on whether AI is applied in the right parts of the process.

Measuring that properly would require a controlled experiment comparing the same researchers working with and without AI assistance. METR says such an experiment would be difficult to organize but would be highly informative. Until then, the expenditure horizon provides evidence about what AI systems can accomplish independently, but much less about how much they accelerate human researchers in practice.