We propose a measure of an AI agent’s optimization ability with an “expenditure horizon.” We give an empirical illustration from the NanoGPT speedrun.
One difficulty in measuring AI’s ability to accelerate AI R&D is accounting for token cost, experiment compute cost and human labor cost. If we can estimate performance as a function of cost for both humans and agents, we can measure the “expenditure horizon” as the point at which those curves cross: the budget at which humans become more cost-effective than AIs. We illustrate our method with data from the NanoGPT speedrun. We first estimate the returns to human effort in optimizing NanoGPT: on the margin, each 1% improvement costs very roughly \$2,500 in human labor. We also report the status of preliminary agentic optimization runs against NanoGPT: after more than \$10K of expenditure we estimate expenditure horizons of \$0-\$3K.
Introduction
A critical question is the degree to which AI will accelerate its own progress.
Over the last year there have been many credible claims that AI has begun accelerating algorithmic progress but it is hard to tell by how much, and what to expect in the future.1 We have many imperfect sources of evidence on AI’s acceleration of AI R&D:
-
    AI R&D benchmarks. There are a few high-quality benchmarks which test agents on their ability to optimize AI training algorithms. However most do not report any human baseline, or only a baseline score at 8 hours or 40 hours. Additionally, the benchmark problems are often not reflective of frontier-level AI R&D. It seems plausible that agents can get good solutions to many “textbook” or “toy” problems without being able to contribute to highly-optimized frontier algorithms.
-
    Researcher-uplift studies. Ideally we would compare researcher productivity with and without AI help, but it is very difficult to run an experiment that estimates the effect accurately (discussion in Becker et al., 2026). We do have some self-reported estimates of productivity improvement: around 4X from the Mythos Preview system card; around 2X from Becker/METR, but both reports recommend a great deal of caution in interpreting these claims. Recent studies find AI significantly increases lines-of-code produced (Anthropic reports that Mythos increased merged lines of code by 8X), but production of code is difficult to map to productivity in R&D.
-
    Qualitative reflections. The Mythos Preview system card describes (1) survey results on ability (4/18 respondents thought Mythos Preview could replace an entry-level researcher with 3 months of scaffolding iteration); and (2) discussion of strengths and weaknesses of Mythos Preview in doing AI R&D. These are useful but again difficult to interpret as quantitative measures of acceleration.
-
    Contributions to frontier optimization problems. There have been many recent reports of autonomous or AI-assisted contributions to frontier optimization problems (some of which are AI R&D problems) e.g. TTT-Discover, AlphaEvolve, and LLM-assisted NanoGPT speedrun contributions. However it is difficult to assess the magnitude of these contributions relative to human effort, and reporting is biased towards successes making the results hard to generalize.
-
    Acceleration in capabilities progress. The evidence above was regarding AI’s contribution to AI R&D inputs; we can additionally look at the trajectory of output, e.g. looking at acceleration in capabilities as measured by METR’s time-horizon or Epoch’s ECI. Estimating how much progress is attributable to AI R&D requires also estimating the contribution of human R&D and training compute (e.g. as done in the Mythos Preview system card). This is intrinsically a lagging metric, only observed after the capabilities are realized in a model.
The method in this note is a novel metric for summarizing an agent’s ability on a optimization problem. It thus can be used to summarize performance on existing AI R&D benchmarks (e.g. RE-bench, MLE-bench), but we also show how to apply it to agentic contribution to certain frontier optimization problems, and we illustrate with NanoGPT.
An agent’s “expenditure horizon” gives a quantitative measure of agentic optimization ability.
We define an agent’s “expenditure horizon” over an optimization problem as the dollar value at which the improvement to the goal metric is equal to the improvement by a human with the same budget.
In the graph below, the green dotted line represents the expected returns to expenditure on human optimization, while the red curve represents the returns to expenditure on optimization by an agent. The expenditure horizon measures the horizontal distance between the starting point and the intersection with the expected human curve.
Measuring an expenditure horizon requires (1) measuring an inference-scaling curve for agent expenditure (the red path); and (2) estimating the local returns to expenditure on human labor (the green line).2
This is a generalization of AI R&D benchmarks which use binary thresholds based on the average human time to achieve a certain score, e.g. RE-bench, MLE-bench, etc. (e.g. as reported in Mythos Preview system card and GPT-5.5 system card). Expenditure horizon differs from these reports in a few ways:
- Reporting a continuous score, instead of a binary score, gives a much more statistically precise measure of capability for each task, i.e. differences between models can be detected with fewer observations.
- Testing against frontier optimization problems, that have already been extensively optimized by humans, makes the outcome more interpretable for real-world usefulness.
- Reporting the scaling curve makes the outcome more relevant to real-world acceleration, e.g. see Noam Brown’s recent exhortation to report scaling curves in benchmarks. Additionally monetary values incorporate task-specific costs of labor and the cost of experiment compute (for frontier AI R&D, the cost of experimental compute is a substantial share of cost).
We discuss limitations in depth below. Many of these limitations apply to any measure of optimization ability, e.g. in typical AI R&D benchmarks. Some important limitations worth highlighting:
- This method only estimates the value of autonomous optimization. In some cases we would expect hybrid optimization (humans assisted by AI) to be a significant improvement on autonomous optimization.
- This method is most useful for problems which show fairly smooth returns to labor. If the returns are lumpy then both human and agent curves are harder to measure.
- The estimated expenditure horizon will likely be shorter on problems for which other agents have already contributed optimizations.
- The estimated expenditure horizon for an agent on a problem will be unrepresentatively short if the lab producing that model has already trained against that problem.
The existence of an intersection assumes that the returns to expenditure on agents diminish more quickly than the returns to expenditure on humans. This seems to be a good description of the current status of agents (they outperform humans at low budgets, but underperform at high budgets, e.g. RE-bench (2024), PaperBench (2025)). However at some point this will no longer hold, after which an agent’s “expenditure horizon” will not be well-defined. If this holds across the frontier AI R&D stack then we will have “automated AI R&D” by many definitions (e.g. those of many labs’ Responsible Scaling Policies, and Ajeya Cotra’s “parity” milestone of AI R&D automation). After this point we can characterize the returns to agentic labor in the same way we measure returns to human expenditure, e.g. a dollar value per 1% efficiency gain, or an elasticity between expenditure and efficiency.
RE-bench (2024) and PaperBench (2025) both document scaling curves for agents and humans (either in time or expenditure), which show an intersection, but they do not calibrate agents’ abilities by their intersection with the human curve.3
The idea of formalizing an agent’s expenditure horizon is partly based on our earlier apple-picking model, although the model need not be taken literally. We find this model useful to answer puzzles in calibrating agent performance against humans. E.g., if an agent can optimize a problem better than a human for \$1000, then why not run the agent twice and get twice the gain? The apple-picking model shows that, if agents are limited in the types of optimizations they can find, then we will expect: (1) agents will outperform humans at small budgets, but fall behind at large budgets; (2) the performance of agents relative to humans will be independent of the cumulative human expenditure on optimization; (3) the performance of agents relative to humans will be dependent on the cumulative agent expenditure on optimization.
We estimate the returns to human labor on NanoGPT.
We give a proof-of-concept showing how to measure expenditure horizon on the NanoGPT speedrun, by comparing human and agentic scaling curves.
We conducted two interviews with prolific NanoGPT contributors. Their responses imply the effort required for an incremental 1 percentage-point optimization is around 16 hours of labor, or \$2,400 (at \$150/hour). We also use an LLM judge to categorize recent contributions to NanoGPT, which estimates a roughly similar return.
These estimates are highly uncertain, especially due to estimating the value of human time — the true cost could be significantly higher or lower. However, we discuss below reasons why we believe the general method remains informative even with over-estimates or under-estimates of the returns to human expenditure.
We show agent expenditure curves and horizons on NanoGPT.
We ran six high-expenditure agentic optimization runs starting from record #78 of the speedrun (March ’26, 85.56 seconds of training time). On reproduction our baseline starts slightly slower which is typically expected on NanoGPT due to noise and hardware differences (Appendix D).
The full returns-to-expenditure curves are shown below, and for each we illustrate the expenditure horizon, shown as the intersection with the estimated returns to human expenditure curve (\$2,500/1%).

A few observations:
- Our harness is likely inefficient. We ran the agents with continuous access to 4 H100 nodes, to use for validation of their optimizations. As a consequence the agents ran many experiments, likely inefficiently, and experiment cost comprised around 70-90% of the cost of most trajectories. We expect a more optimized harness would significantly lower cost for a given optimization. However we can see from the curves that shifting the expenditure curves horizontally would not dramatically change the expenditure horizon, and it appears unlikely to increase the maximum speedup achieved.
- The raw trajectories overstate progress. The trajectories show the cumulative best score. But for each run we also re-validate their trajectories. The overstatement is due to statistical noise. For the best performing models we further think only ~70% of contributions are mergeable as per the maintainer’s judgement.
- GPT-5 and Opus-4.1 only chase noise. The raw trajectories from GPT-5 and Opus-4.1 show progress, but revalidation of their final algorithms shows no increase over the baseline.
- Some models show significant improvements. GPT-5.5 and Opus-4.8 show significant improvements over the baseline, and we can quantify those with significant expenditure horizons, as shown above.
- The big-picture improvements are still modest. Although these models have expenditure horizons in the thousands of dollars, they are small relative to the overall expenditure on human labor. Thus this evidence implies that autonomous optimization does not have dramatic effects on AI R&D progress on NanoGPT (though it is still possible that agents could dramatically augment human progress, AKA hybrid optimization).
Theory
In this section we give a more detailed discussion of a variety of methodological issues related to expenditure horizon.
We first note that this method is applicable to constrained optimization problems. The progress of AI R&D is often characterized as optimization efficiency, and many benchmarks use this as a scoring rule (RE-bench, MLE-bench), but there are of course many parts of AI R&D that cannot be easily interpreted as maximizing a pre-existing quantitative metric.
Relation to Time Horizon.
The basic building block of the Time Horizon methodology introduced in Kwa et al., 2025 is measuring whether AI agents succeed or fail on tasks which take humans different amounts of time. Two limitations of this are:
- This uses binary pass-fail scoring, which throws away information if we have richer feedback about task success (e.g. a range of possible grades, or a continuous score like “accuracy”)
- It doesn’t fully specify a budget or constraints for tokens or other resources (e.g. compute for running experiments, calendar time, number of queries to evaluation function) should be used for AI agents or humans. This is less important if agent performance plateaus well below human cost, but becomes important if the agent is still making progress at token spend similar to human labor spend, or if other resources are important bottlenecks.
If a task has approximately continuous scoring, and we can measure or predict human and agent score as a function of some kind of cost, the expenditure horizon addresses the two limitations above: we get more information per agent run than using a binary score cutoff, and it’s clearer how to handle resource budgets for humans and agents.
For example, we can define the “compute and labor” expenditure horizon by denominating experiment compute usage, agent token usage, and human labor cost in dollars, plotting cost-performance curves for agent and human, seeing where the curves cross.4
We expect costs like “experiment compute” or “queries to evaluation function” to be important factors in determining how much AI agents can accelerate AI R&D.
Which optimization problems to use.
We wish to run the agent on a problem (a well-defined constrained optimization problem) and a state (a starting point on that problem, e.g. an algorithm). An ideal problem and state would satisfy the following criteria. We believe the NanoGPT speedrun satisfies most of these:
- The problem is similar to frontier AI R&D. The most important criterion is that this problem reflects typical frontier AI R&D. It is important to keep in mind that all the following additional criteria narrow the scope for the purposes of tractability, and by simplifying the problem of AI R&D they could overstate the usefulness of agents. NanoGPT is notably similar to frontier pretraining algorithms, e.g. the Muon optimizer was first invented for NanoGPT, but has become widely adopted since, although NanoGPT’s scale is far far smaller than any frontier pretraining algorithm.
- The outcome metric is well-defined. In some cases optimization can involve tradeoffs among multiple variables. If the tradeoff among those variables is not well-defined, then it becomes much harder to judge returns to expenditure with a single variable.
- Progress is regular. If historical progress on an optimization problem consists of irregular advances, e.g. occasional breakthroughs after months of effort, then it is much harder to estimate the returns to both human and agent effort. Data shown below indicates NanoGPT progress has been fairly regular over the last year. It is also notable that, by some measures, NanoGPT has had a 700X efficiency improvement since GPT-2, which is reassuring that there are still many more potential efficiencies.
- We have data on the return to human effort. To calibrate the returns to agent effort we need some data on human effort already put in on this problem, and the returns to that effort. Ideally we measure the market price of that effort, reflecting the wage of the people who would ordinarily be working on this problem.
- Progress is cheap to verify. If progress on a problem is expensive to verify then it is difficult to measure the returns to expenditure. NanoGPT progress is somewhat cheap to verify: frontier training time is less than 2 minutes, though training-time is noisy so validation requires averaging many runs. Some AI R&D problems are expensive to verify, e.g. the performance of pretraining algorithms on frontier models may not be accurately observable until after millions of dollars of training compute. Nevertheless we expect some frontier AI R&D to be cheap to verify, e.g. efficiencies in post-training, elicitation, or inference; others have cheap-to-verify proxies, e.g. small derisks for pretraining.
- The starting state does not already reflect significant agentic optimization. If the agent starts from a state that already reflects work done by prior agents then we expect lower returns to additional agentic labor. Intuitively, starting from a state that already has been optimized by another agent is equivalent to starting part-way along an existing agent scaling curve.
- The agent does not know about subsequent states. If an agent is aware of contributions to the problem subsequent to the starting state (either through training or through internet access), then the model’s progress would overstate its capabilities on true frontier optimization problems.
- The agent has not been trained on this problem. If the model’s developer already spent a significant amount of compute post-training against this specific problem, then we would expect the agent to perform very well in the short-run, over-stating the true returns to expenditure on agentic labor at the frontier. Unfortunately OpenAI and Anthropic haven’t publicly disclosed whether they train on NanoGPT, so we can’t be sure here. Ideally we would measure expenditure horizons across a few different hard optimization problems.
Criteria #6 and #7 are in tension: we wish to find a starting checkpoint that’s not too recent (so it doesn’t incorporate substantial agentic optimization), and not too old (so the model isn’t already trained on subsequent states).
Calibrating to money vs calibrating to time.
We believe it is most informative to measure agent and human performance with respect to expenditure (money), inclusive of expenditure on experiments. Returns to expenditure directly measure the economically-relevant variables for an AI R&D lab that is choosing between spending on human vs agentic labor.
However it can also be worth reporting the time-equivalent, by converting the dollar equivalent back to human time using the estimated wage. Notably this is somewhat awkward when expenditure includes a large share on experiments (the experiment cost would be converted to hour-equivalents).5
Estimating human expenditure curves.
The method requires measuring the local returns to labor. The simplest way is to ask people roughly the existing returns to optimization, e.g. the total budget on salaries and compute, and expected returns.
For public optimization problems there already exists a significant literature documenting the returns to human effort across a number of optimization problems, typically summarized with an elasticity r between cumulative effort and efficiency progress, e.g. Erdil et al. (2024).
Bias in human expenditure curves.
Estimating the cost of human inputs can be difficult, and our estimates for expenditure on NanoGPT could be significant under-estimates or over-estimates. Nevertheless we believe our method is fairly robust for two reasons:
-
    Relative expenditure horizons across models are likely to be preserved if we scale the human cost of progress by a fixed ratio. This would be true if, for example, all agents have the same elasticity in returns to expenditure (i.e. the same proportional increase in efficiency for a proportional increase in expenditure).
-
    Our method assumes the returns to agent expenditure diminish more quickly than the returns to human expenditure. This implies rescaling the human curve will have a relatively larger effect on the expenditure intersection than on the efficiency intersection, so we expect the efficiency gain from using agents to be relatively stable to rescalings.
Estimating hybrid optimization ability.
This note has compared human-only vs agent-only optimization, however we are often interested in hybrid optimization ability, where the human is helped by an agent. We illustrate a hypothetical hybrid scaling curve below:
If humans make efficient choices in the use of LLMs then we expect the hybrid curve to dominate both the human and agent curves. However there is also evidence that hybrid performance can be worse than human-only performance, e.g. Becker et al. (2025).
Ideally we would estimate the hybrid curve by an experiment, comparing humans with and without AI assistance. In practice this is difficult to organize, but it would be a very informative experiment.
An alternative approach is to measure the cost saving.
Given the agent and human expenditure curves, another metric of interest is the point at which the two curves are tangent, illustrated with the blue point below. From this point we can estimate the “cost saving” due to agentic ability.
The cost saving metric reflects the theoretical economic value of an agent somewhat better than expenditure horizon, but we find expenditure horizon easier to describe and statistically easier to estimate.