METR proposes 'expenditure horizon' to price agents vs humans
METR's new metric finds the budget at which AI research agents stop being cheaper than people — currently in the low four figures for only the newest models.
Measuring capability in dollars
The evaluation group METR published a new metric on July 27 that reframes agent capability as a budget question. Rather than asking whether a model can complete a task, the expenditure horizon asks how much money you can usefully spend on an agent before a human researcher becomes the cheaper way to buy the same improvement. Below that threshold the agent wins on cost; above it, you should hire.
METR derived the figure from the NanoGPT speedrun, a community benchmark in which contributors compete to reduce training time for a small language model. Human contributors, by METR's accounting, spend roughly $2,500 of effort per one-percent speedup. Against that baseline, only the newest frontier systems — GPT-5.5 and Opus-4.8 in METR's runs — register a meaningful expenditure horizon at all, and it lands in the low four figures. Total human effort behind the speedrun is estimated at around $250,000; autonomous AI contributions moved the needle by a small fraction of that.
A qualitative finding sits alongside the numbers: roughly 70% of agent-generated ideas were technically integrable, but most were unoriginal — recombinations of known techniques rather than the novel moves that produce the largest gains.
The stated caveat
METR measured fully autonomous optimisation only, which is not how research is actually done. The group notes that human-plus-AI configurations sometimes underperform humans working alone, and that its metric says nothing about which collaboration patterns work.
Why it matters
Agent capability has been argued almost entirely in benchmark percentages, which give buyers no way to decide how much compute a task deserves. A dollar-denominated horizon converts frontier progress into a procurement number and makes it trackable across releases — if the horizon climbs from four figures to five, the automation frontier has genuinely moved regardless of what any eval score says. The current reading is sobering for R&D automation claims: on a real, competitively optimised research task, autonomous agents remain cheaper than humans only within a budget most labs would spend in an afternoon.