Artificial Intelligence (AI)

When More Tokens Make AI Output Worse, and When They Make It Better

The relationship between AI token usage and output quality is not a straight line, and it has been measured repeatedly: compute-optimal test-time scaling beat a model fourteen times larger by up to twenty-eight percent on easy and medium questions while losing thirty-seven percent on hard ones, inverse scaling in test-time compute shows longer reasoning degrading accuracy across four separate task families with failure modes that differ by model family, models under extended thinking budgets abandon correct answers they had already generated earlier in the chain, all eighteen models in Chroma's context rot study degraded as input length grew even on deliberately trivial tasks with a three hundred token focused prompt beating a hundred and thirteen thousand token full context, and METR's randomized trial found experienced developers were nineteen percent slower with AI tools while believing they had been twenty percent faster.

Somewhere in your organization there is probably a number that counts how much AI you are using. A dashboard, a seat report, a line in a board deck. The implied theory is that the number going up is good.

The measured relationship between ai token usage and output quality is not a straight line. It has been studied more carefully than the year’s coverage suggests, and the findings are mostly a case against spending more, with one well-defined exception.

This is the reference version of that evidence. We have covered what tokens actually are and what input and output tokens buy you. This is the question those two set up: does spending more of them produce better work.

Where the word came from, and why it is already over

The shorthand for spending as much as possible is tokenmaxxing, and its lifecycle is short enough to summarize in a paragraph.

Kevin Roose wrote it up in The New York Times on March 20, 2026, describing a status game in which employees at large technology companies competed on internal leaderboards to show how much AI they were using. Nobody has documented who coined it. Every account says it circulated in group chats first and Roose gave it a byline.

It peaked in April and May. By May, Amazon had shut down its internal token leaderboard with the instruction not to use AI just to use AI, after employees acknowledged inflating their numbers with irrelevant tasks. By May 28, Fortune had published the obituary. Forbes followed on June 2 with the replacement term, valuemaxxing, and O’Reilly ran "The End of Tokenmaxxing" on June 30. Google Cloud’s 2026 AI ROI report is subtitled "From tokenmaxxing to ROI" and reports four times more search interest in token efficiency than in tokenmaxxing or token leaderboards since January.

So the word is six months past its peak and the vendors have already sold everyone the sequel.

We are not writing the explainer. We are writing the part almost none of that coverage included, which is what anyone actually measured. And the first thing worth saying about that literature is this:

None of the research uses the word. Every study cited below was conducted about test-time compute, reasoning budgets or context length. Not one of them is about tokenmaxxing, and several predate the term. The term has essentially no measured evidence of its own, in either direction. It should not be allowed to borrow the credibility of work that was not done about it.

Why a bad metric was so attractive

Worth pausing on, because the specific word is gone and the structure that produced it is not.

The clearest description of what happened came four days after the Times piece, when Duncan Grazier called tokenmaxxing lines-of-code thinking for the agentic era. That comparison earns its keep. Counting lines of code failed for reasons everybody can recite: it rewarded verbosity, punished deletion, and measured the byproduct of the work instead of the work. Every one of those objections transfers intact.

The reason it came back anyway is that the alternative is genuinely hard. Token consumption has three properties that make it irresistible to anyone who has to report upward. It is available immediately, with no instrumentation to build. It is precise to the digit, which reads as rigor. And it is not arguable, because nobody can claim the number is wrong.

A measure of whether the AI actually helped has none of those properties. It takes a quarter to define, it requires agreeing in advance what the work was supposed to produce, and every figure it yields is contestable. Faced with a board or a client asking how the AI investment is going in November, the precise meaningless number wins on availability alone.

This is ordinary Goodhart’s law, and the organizations we work with are not immune to it because they are small. A fifteen-person nonprofit does not run an internal leaderboard. It does get asked by its board whether the new AI tooling is worth the line item, and the easiest thing to put on the slide is the seat count and the usage graph. That is the same mistake at a different scale, arriving about a year later.

The rest of this piece is the material you would want before that meeting.

What the spending numbers are, and what they are not

The figures that circulated during the peak are worth repeating once, because they set the scale, and then setting aside, because of what they are.

Meta reportedly ran 60.2 trillion tokens in thirty days, roughly 900 million dollars at list API rates, with a single top individual at 281 billion. One OpenAI engineer reportedly used 210 billion tokens in a week. One Claude Code user reportedly ran a 150,000 dollar bill in a month. Uber reportedly exhausted its entire 2026 AI budget by the end of April. Salesforce set internal minimums of 100 dollars a month on Claude Code and 70 on Cursor. In March, Jensen Huang suggested a developer on a 500,000 dollar salary should be spending something like 250,000 a year on tokens.

Every one of those is an input. None is an outcome.

That distinction is the whole problem, and it is not a new one. We made the same argument about vendor benchmarks and leaderboards in August: a number that is easy to produce and hard to dispute will get optimized whether or not it measures anything.

The nearest thing to a broad outcome measurement is Google Cloud’s 2026 AI ROI report, which surveyed 2,403 executives through National Research Group. It found 84 percent reporting increasing financial returns from AI, but only 26 percent describing those returns as accelerating. That gap between returns existing and returns compounding is the honest version of where most organizations sit.

It is also a vendor survey with an obvious narrative interest, published by a company selling the successor framing. Directionally useful, not independent, and we would not build an argument on it alone.

There are outcome-flavored numbers in circulation, and we are going to decline to use them. A figure of one company finding its heaviest AI users were twice as productive at ten times the token cost has been repeated widely. So have claims of an 861 percent increase in deleted lines of code and a 39 percent increase in errors. We could not trace any of them to a published methodology. If you have seen those numbers, that is where they come from, which is nowhere.

The one place more tokens is measured to help

There is a real result here, it is good work, and it is more conditional than the summaries suggest.

Snell and colleagues, in "Scaling LLM Test-Time Compute Optimally Can Be More Effective Than Scaling Model Parameters" at ICLR 2025, asked whether letting a model think longer at inference beats training a larger model. Sometimes it does.

Allocating test-time compute adaptively, rather than uniformly, matched a best-of-N baseline using roughly four times less compute, 32 generations against 128. Compared against a model about fourteen times larger on a matched compute budget, extra inference compute beat extra parameters by up to 27.8 percent on easy and medium questions at low inference load.

The mechanism behind that first number is worth separating from the headline, because it is the transferable part. The gain did not come from spending more. It came from spending adaptively, allocating effort according to how hard a given problem appeared, rather than applying the same budget to everything. Uniform generosity is what the four-times-less-compute figure was measured against, and it lost.

Then the part that gets left out. On hard questions the same comparison ran 37.2 percent worse in relative accuracy. Pretraining wins there. The benefit is not a property of thinking longer. It is a property of thinking longer about problems that are within reach.

That conditionality is the single most useful thing in this entire literature, and it maps onto ordinary work cleanly. Turning up reasoning effort on a task the model can nearly do is a good trade. Turning it up on a task the model fundamentally cannot do buys you a longer, more confident wrong answer.

Four task families where longer reasoning makes it worse

The strongest counter-evidence is peer reviewed, and it is specific about where the failure happens.

"Inverse Scaling in Test-Time Compute", by Ethan Perez and thirteen co-authors, was published in TMLR in December 2025 with a Featured Certification. It constructs tasks where extending the reasoning chain makes accuracy go down, and it finds four families where that happens:

  • Counting with distractors. Simple counting problems with irrelevant information present.
  • Regression with spurious features. Prediction tasks where a misleading correlation is available.
  • Deduction with constraint tracking. Logic problems requiring multiple constraints held at once.
  • AI risk evaluation. Assessment tasks in that domain.

The detail that matters most for practice is that the failure modes differ by model family. Claude models became increasingly distracted by irrelevant context as reasoning length grew. OpenAI’s o-series resisted distractors but overfit to the framing of the problem. Same intervention, opposite weaknesses.

If you have standardized on one vendor and tuned your prompts around its behavior, that tuning does not transfer. A prompt that works because it aggressively strips irrelevant context is compensating for one failure mode. A prompt that works because it restates the problem three ways is compensating for the other.

Models throw away answers they already had

A second result, from April 2026, explains part of the mechanism and is the one most likely to change how you configure a tool.

"When More Thinking Hurts" finds that returns diminish substantially at higher reasoning budgets, that optimal thinking length varies by problem difficulty, and that models abandon correct answers they generated earlier in the chain. The paper’s conclusion is that comparable accuracy is available at substantially lower cost by limiting reasoning length rather than extending it.

Read that middle finding slowly. The model had it. Then it kept going, and talked itself out of it.

That is not a subtle statistical effect to be weighed against a countervailing benefit. It is a direct argument that the maximum setting is the wrong setting, and that the right setting is a function of the task rather than a default you turn up once and forget.

Context rot: every model tested got worse as input grew

The result with the most direct consequence for anyone building on a CMS comes from Chroma’s context rot research, published July 14, 2025 by Kelly Hong, Anton Troynikov and Jeff Huber.

They tested eighteen models, including GPT-4.1, Claude 4, Gemini 2.5 and Qwen3. Every one of the eighteen degraded as input length grew. That held even on tasks deliberately designed to be trivial, where length should not have mattered at all.

Two specifics are worth carrying around:

On LongMemEval, models did markedly better with a focused prompt of roughly 300 tokens than with the full 113,000-token context containing the same answer. The information was present in both. The short one worked better.

Shuffled haystacks outperformed coherently structured ones. Randomizing the surrounding material produced better retrieval than presenting it in a logical order, which is the opposite of what almost every retrieval pipeline is built to do.

This lands directly on retrieval-augmented generation work. The instinct when a RAG answer is wrong is to retrieve more chunks, widen the window and give the model more to work with. The measurement says that instinct is backwards past a point that arrives earlier than anyone expects. A tighter retrieval step that returns three good chunks will beat a generous one that returns thirty, and it will cost a fraction as much.

The shuffled-haystack result deserves a second look before anyone dismisses it as a curiosity. Retrieval pipelines are almost universally built to assemble context coherently: related passages grouped, chronology preserved, sections in document order. The measurement says that coherent assembly performed worse than randomized assembly on the same material. Nobody should rewrite a pipeline on one finding, and nobody should assume their careful ordering is helping either. It is a cheap thing to test and an expensive assumption to leave untested.

It also means that large context windows are a capability, not an instruction. Being able to pass 200,000 tokens does not mean you should.

The measurement problem underneath all of it

One more study, and it is the one that should make anyone uneasy about self-reported AI productivity of any kind.

METR ran a randomized controlled trial with experienced open-source developers working on their own repositories. The developers were 19 percent slower when using AI tools. They estimated afterward that they had been 20 percent faster.

A thirty-nine point gap between what happened and what it felt like.

That number does more damage to the case for usage metrics than any of the reasoning research does, because it undermines the instrument. If practitioners cannot feel the difference between a nineteen percent slowdown and a twenty percent speedup, then survey data about AI productivity is not measuring productivity. And a token counter is worse than a survey, because it does not even ask.

We are not claiming this generalizes to every task or every tool. It was a specific population, experienced developers on mature codebases they knew well, which is close to the best case for human speed and a hard case for a model. It also predates the current generation of tools, which is a fair objection and a limited one. It moves the size of the effect; it does not restore the instrument. A study finding that developers misjudged their own speed by thirty-nine points is evidence about human estimation under this class of tool, and no model release fixes that.

The finding that generalizes is narrower and more useful: perceived productivity and measured productivity came apart badly, in the direction of overestimating the tool.

Three failure modes, not one

Read together, the studies above are not four versions of the same finding. They describe distinct ways the same intervention goes wrong, and they call for different responses.

Attention dilution. The context rot result is about input. As the surrounding material grows, the model’s handling of the part that matters degrades, and it degrades even when the task is trivial and the answer is present. Eighteen models, no exceptions. The response is to send less.

Chain drift. The "When More Thinking Hurts" result is about output. Given a longer budget, a model that had already produced a correct answer continues, revises, and lands somewhere worse. The response is to cap the budget rather than raise it.

Systematic distraction and framing lock-in. The inverse scaling result is about what a longer chain does to a model’s priorities, and it splits by vendor. Longer reasoning made Claude models progressively more susceptible to irrelevant context. It made the o-series more committed to whatever framing the prompt supplied. The response differs by which one you are on, which is the uncomfortable part for any team maintaining one prompt library across two providers.

None of these is a bug awaiting a patch, and it is worth being careful about that. They are properties measured across current frontier models by independent groups using different methods, and at least two of the three get worse as models get more capable at long-form reasoning, not better. Planning around them is the job.

What the three have in common is the shape of the curve. There is a region where additional spend buys accuracy, a peak, and a region past it where additional spend buys confident error. Every result here is a measurement of where that peak sits for some class of task. None of them found that the curve keeps rising.

What this means on a real project

Translating the above into decisions, for the kind of work we do.

Set reasoning effort per task, not per project. The Snell result says more inference compute helps on problems near the edge of a model’s ability and hurts on problems past it. The "When More Thinking Hurts" result says the optimal length varies by difficulty. Both point the same way: a single global setting turned up to maximum is wrong for most of what passes through it.

Retrieve less. If you are building RAG on a client site, the context rot finding says your retrieval step should be tighter than feels comfortable. Fewer chunks, better selected. Test a deliberately narrow configuration against your generous one before assuming more is better, because eighteen out of eighteen models said otherwise.

Do not paste the whole document. The 300-token prompt beating the 113,000-token context is the version of this that applies to everyday use, not just to pipelines. Extracting the relevant section by hand is not a workaround for a limitation. It is the better input.

Know which failure mode your vendor has. If you are on Claude, irrelevant context in a long chain is your risk, so prune inputs. If you are on the o-series, framing lock-in is your risk, so vary how you state the problem. Prompt libraries tuned for one do not port to the other.

Measure the short configuration against the long one, on your own work. Every number in this piece came from someone else’s task distribution. The peak of the curve sits in a different place for summarizing meeting notes than for refactoring a theme. The cheap version of this is to run both configurations on twenty real items from your actual backlog and count how many came back needing rework. That is a morning’s work and it beats every benchmark in this article for your purposes, because it is about your work.

Write down what a good outcome is before you measure anything. This is the part that has nothing to do with models. A token count is available on day one and means nothing. A defensible measure usually takes a quarter to define and requires agreeing what the work was supposed to produce. That is why the token count wins, and it is why it should not.

Is valuemaxxing any better

The replacement framing arrived fast, which is itself worth noticing. Forbes had the successor term in print five days after Fortune ran the obituary, and IBM, Google Cloud and most of the analyst layer were using it within the month.

The substance of valuemaxxing, stripped of the branding, is: measure outcomes rather than consumption. That is correct. It is also what anyone who thought about it for ten minutes would have said in March, and it is not a methodology. None of the material we read attached a definition of value, a way to compute it, or an instrument for collecting it.

So treat it as a course correction rather than an answer. The useful part is that the vendor layer has publicly abandoned usage as a success metric, which means you can now push back on a usage-based dashboard by citing the companies that sold it to you. The unhelpful part is that the replacement dashboard will be measuring something equally convenient unless somebody does the hard definitional work, and the hard definitional work is still yours.

The thing to watch for is a value metric that is just a usage metric with a multiplier on it. A slide showing tokens consumed times an assumed hourly rate saved is not an outcome measure. It is the same number wearing a jacket.

What to measure instead

Short list, and none of these is novel. They are just harder than counting tokens, which is the entire reason they are not what anyone is counting.

  • Task completion rate. How often does the AI-assisted attempt finish without a person redoing it.
  • Rework. How much of what it produced came back. The deleted-lines figure circulating without a methodology was reaching for this, and reaching for it is not the same as measuring it.
  • Time to a reviewed, merged, shipped result, not time to a first draft. The METR result exists precisely in the gap between those two.
  • Error rate in production, attributed honestly.
  • Cost per completed unit of work, which is the only place token spend belongs: as a denominator, never as a numerator.

We looked at the budgeting side of this in managing AI investments in the agentic era. The short version holds up: useful work per dollar is the only framing that survives contact with the research.

What we could not establish

Four honest gaps, because a reference piece that hides them is not a reference.

Who coined tokenmaxxing. Every account traces it to Silicon Valley group chats before March 2026. We found no attestation earlier than the Times piece despite looking for one. Roose popularized it; the coiner is undocumented.

Whether the practice ever worked. There is no study of tokenmaxxing as such, in either direction. The absence of evidence that it helped is not evidence that it hurt. What exists is a body of adjacent research, conducted for other reasons, that mostly points against maximizing.

How far the METR result generalizes. One population, experienced open-source developers on repositories they knew well, which is close to the best case for unaided human speed. We have not seen an equivalent trial on junior developers, on unfamiliar codebases, or on non-engineering work, and we would read one with interest.

Where the peak of the curve sits for your work. Nobody can tell you this from the outside. Every measurement here used a task distribution that is not yours. That is why the practical recommendation ends with running your own comparison rather than adopting a number from this article.

Frequently Asked Questions

What is tokenmaxxing?

A term for treating AI token consumption as a productivity metric, and the behavior that follows from that, including internal leaderboards and inflated usage. Kevin Roose popularized it in The New York Times in March 2026. Nobody has documented who coined it.

Is tokenmaxxing still a thing?

Not really. It peaked in April and May 2026. Amazon shut down its internal token leaderboard in May, Fortune ran the obituary on May 28, and by June the industry framing had shifted to token efficiency and valuemaxxing. Google Cloud reports four times more search interest in token efficiency than in tokenmaxxing since January 2026.

Does spending more tokens improve AI output?

Sometimes, and the condition is specific. Compute-optimal test-time scaling beat a fourteen times larger model by up to 27.8 percent on easy and medium questions, and ran 37.2 percent worse in relative accuracy on hard ones. More tokens help on problems near the edge of a model’s ability and hurt on problems beyond it.

What is inverse scaling in test-time compute?

A measured effect in which extending a model’s reasoning chain makes accuracy worse. The TMLR paper documenting it identifies four task families where this happens: counting with distractors, regression with spurious features, deduction with constraint tracking, and AI risk evaluation.

Do different models fail in different ways?

Yes, and this is the practical finding. Under extended reasoning, Claude models became increasingly distracted by irrelevant context, while OpenAI’s o-series resisted distractors but overfit to the framing of the problem. Prompting strategies tuned for one do not transfer to the other.

What is context rot?

Degradation in model performance as input length grows. Chroma tested eighteen models and found it in all eighteen, including on tasks deliberately designed to be trivial. A focused 300-token prompt outperformed a 113,000-token context containing the same answer.

Should I use the biggest context window available?

Not by default. A large context window is a capability, not an instruction. The measured evidence says performance degrades with length across every model tested, so passing more because you can is a cost with a penalty attached.

How does this change how I build RAG?

Retrieve less and select better. The instinct to widen the window when answers are wrong runs against the measurement. Test a deliberately narrow retrieval configuration against your current one rather than assuming more context helps.

Can I trust my team’s sense of whether AI is helping?

Cautiously. METR’s randomized trial found experienced developers were 19 percent slower with AI tools while estimating they had been 20 percent faster. Perceived and measured productivity came apart by 39 points, in the direction of overestimating the tool.

Does any of the research actually study tokenmaxxing?

No. Every study cited here was conducted about test-time compute, reasoning budgets or context length, and several predate the term. The term itself has no measured evidence behind it in either direction.

What should I measure instead of tokens?

Task completion rate, rework, time to a reviewed and shipped result rather than to a first draft, production error rate, and cost per completed unit of work. Token spend belongs in that last one as a denominator and nowhere else.

Is there a case where a token count is a useful number?

As a cost input for forecasting and capacity planning, yes. As a measure of output, effort or productivity, no. The distinction is whether it sits under the line or over it.

Digital Matters

Artificial Intelligence (AI) Desk