When AI Thinks Longer: Test-Time Compute and the Reasoning Bet

Some new AI models spend extra time on a hard question before they answer.  One long ramble often gets worse, not better.  A few shorter independent tries, then picking the answer most of them agree on, usually works better.

You will see that bet on models such as OpenAI’s o1 series and DeepSeek-R1.  They spend more computer work at the moment of the answer — a longer private scratch pad, extra attempts, sometimes a checker — to get more hard math, code, and science questions right.  The “think longer” bet is real.  It is also less simple than a rule that more words always help.  Lab work from 2025–26 shows that stretching one attempt can tip into overthinking, while a few shorter tries side by side, then keeping the answer most of them share, often goes further on the same budget.

Rows of high-performance computing servers in a data center, lit by blue aisle lights
Inference is where the new reasoning bill shows up — GPU aisles keep humming after training ends. (Unsplash)

As an Amazon Associate I earn from qualifying purchases. If you buy through links on this page, I may earn a commission at no extra cost to you.

What “test-time compute” actually means

For about a decade, the usual story for large language models was training, not the moment someone asks a question.  More data, more parameters (the internal settings a model learns), and more computing while the model is built.  Papers sometimes count that training work in FLOPs, short for floating-point operations, the tiny math steps a chip performs.  The plain version is that “smarter” mostly meant “trained harder.”

Reasoning models add a second dial.  Test-time compute is extra computing spent while the model is answering you, after training is already finished.  Before it commits, it may allow itself a larger token budget, which is only a cap on how much text it may generate, including hidden scratch work.  That allowance can feed a long private chain of thought (the step-by-step scratch pad written before the visible answer), several candidate answers, search-like revision (going back and reworking one attempt), or a separate check that the final answer hangs together.

OpenAI’s September 2024 introduction of o1-preview said it directly: models designed to spend more time thinking before they respond, with stronger results on science, coding, and math than earlier chat models on the same hard tests, which labs call evals.  In the companion note “Learning to reason with LLMs,” OpenAI reported that o1’s performance improved with more reinforcement-learning compute during training — a training style that rewards better reasoning — and with more time spent thinking at test time.  Two different knobs, not one.

By December 2024, the production o1 API release added a practical control: a reasoning_effort parameter so developers could trade latency (how long a person waits) and cost against thoroughness.  OpenAI said that snapshot used fewer reasoning tokens — chunks of hidden scratch text — on average than o1-preview for a given request.  The product lesson landed early.  “Think longer” is not only a research curve.  It is a meter that customers eventually see on an invoice.

DeepSeek-R1 and the open reasoning wave

In January 2025, DeepSeek published DeepSeek-R1 (arXiv:2501.12948).  It is a large-scale reinforcement-learning approach that encourages reasoning behaviors such as reflection (looking back at its own steps), verification (checking the work), and strategy shifts (trying another approach), without depending on human-written chain-of-thought demos — example scratch pads — for the whole pipeline.  The paper and the GitHub release put the model weights and a detailed write-up in public view, which mattered for the research conversation even more than any single benchmark row.

DeepSeek’s materials contrast R1’s adaptive token use with fixed majority-vote or search recipes: fewer tokens on easy items, more on hard ones, instead of the same script every time.  A majority vote means taking several answers and keeping the one most of them agree on.  They also note a familiar failure mode — overthinking on simpler questions, where the model keeps reasoning past the point of usefulness.  That admission is useful.  It matches what independent 2025 papers found when they tried to stretch one chain of thought forever.

Chalkboard covered with mathematical formulas and diagrams
Hard math and coding benchmarks are where longer reasoning traces earn their keep — and where overthinking shows up first. (Unsplash)

When thinking longer stops helping

A tidy slogan — “more thinking equals better answers” — does not survive contact with careful measurement.  Several 2025 studies looked at o1-like and R1-like models and asked whether extending a single thinking trace is the best use of a fixed inference budget (the computing spent on this one answer).  A thinking trace is one continuous scratch pad.

One line of work, summarized in “Does Thinking More Always Help? Mirage of Test-Time Scaling in Reasoning Models” (arXiv:2506.04210), reports a non-monotonic pattern, meaning the score does not keep climbing.  Accuracy can rise with extra thinking steps, then plateau or fall as models revise, second-guess, or abandon an already-correct path.  The authors argue that some apparent gains from “wait, rethink” prompting are a mirage tied to uncertainty and to how the test scores behave.  They are not a clean proof that sequential revision — staying on one answer and rewriting it again and again — is the right scaling law, a reliable rule for how extra computing turns into better answers.

A related February 2025 study, “Revisiting the Test-Time Scaling of o1-like Models” (arXiv:2502.12215), compared sequential revision with parallel sampling under similar token budgets (a similar allowance of generated text).  Parallel sampling means several separate tries at the same time, rather than one try that keeps editing itself.  Parallel coverage improved faster.  Sequential revision helped only up to a point, and it carried a higher attention cost over long contexts: extra computing just to keep a very long write-up in mind.  Their “Shortest Majority Vote” variant preferred clusters that were both popular and short — a practical hint that the first clean answer is often better than the longest one.

“Don’t Overthink it” (arXiv:2505.17813) pushes the same theme harder.  Within a question, shorter sampled chains were substantially more likely to be correct than the longest chain for that same question.  Their short-m@k method — the lab name for starting several answers at once, stopping when the first m finish, and voting among those shorter ones — aims for majority-vote quality with less wall-clock time and fewer thinking tokens.  Here m is just how many early finishers they wait for, wall-clock time is time on an ordinary clock, and thinking tokens are text chunks spent on scratch work.

None of this says reasoning models are a fad.  It says the shape of the computing matters.  Serial “keep talking” is not the same as a few independent attempts and the answer most of them share.

Several tries, a shared answer, and a checker

Older recipes still matter.  Sample N answers — generate N separate tries — and take the majority, or sample N and let a reward model or a verifier pick a winner.  That second recipe is called best-of-N.  A reward model is another model trained to score which try looks better.  A verifier is a checker: it tests whether an answer holds, not only whether it sounds fluent.  Those ideas predate o1, but reasoning models made them newly expensive — and newly valuable — because each sample can be a long, costly chain of thought.

The 2025 “mirage” paper’s parallel-thinking alternative is straightforward.  Split the same token budget across independent paths and vote.  In their controlled comparisons, that beat stretching one path with “wait / think more” style prompts by a wide margin under a fixed budget.  You do not need to treat the exact percentage lifts as universal laws.  The directional lesson is enough for product teams.  A variety of attempts often buys more than a longer monologue.

Verifiers add another layer when the subject allows a real check — unit tests for code, a symbolic checker for math, or a smaller judge model.  OpenAI’s own January 2025 note on trading inference-time compute for adversarial robustness — holding up when an input is built to trick the model — also frames longer thinking as a resource that can improve behavior under attack in some settings, while acknowledging cases where more computing does not help.  Again: the budget is a tool, not a magic wand.

Cost, waiting, and what products should ship

For people building products, the interesting question is not “does test-time compute work?” It is “when is the bill worth it?”

  • Hard, verifiable tasks. Contest math, tricky coding, and multi-step scientific questions — domains where a wrong answer is obvious after the fact — are where extra computing at answer time has repeatedly shown up in vendor evals (scored lab tests) and in open research.
  • Easy or fuzzy tasks. Short factual lookups, tone edits, and open-ended brainstorming often do not need a marathon chain of thought.  Overthinking here burns latency (the wait a person feels) and money for little gain.
  • Interactive products. Users feel seconds.  A control in the style of reasoning_effort, a cheap first pass with a “think harder” button, or an async “deep solve” job — slow work that runs in the background — is usually kinder than making every chat turn wait for a long private monologue.
  • Batch and agent workloads. Overnight evals, code agents with sandboxes, and research pipelines can afford more parallel samples (side-by-side tries).  A sandbox is a contained place to run code.  A majority vote, and early-stop tricks in the style of short-m@k, fit those settings better than a single ultra-long trace.

There is also an infrastructure angle.  Training still dominates the biggest capital spends, but inference — the computing that runs every time someone asks — is where the perpetual operating cost lives.  If every hard query multiplies tokens by ten or a hundred, capacity planning and pricing have to change — even when accuracy charts look great in a paper.

The bottom line

Test-time compute is a real second axis of progress for language models.  OpenAI’s o1 work and DeepSeek’s R1 release made “spend more to think” a mainstream product and research story.  The follow-up science is just as important.  Longer is not automatically smarter.  Parallel attempts, a majority vote, shorter preferred chains, and verifiers that fit the subject often stretch a fixed budget better than one endless revision spiral.

If you are shipping a reasoning feature in 2026, treat the thinking budget like any other scarce resource.  Put the heavy dial on hard, checkable work.  Keep a fast path for everything else.  And measure accuracy against wall-clock time and dollars — not against how impressive the hidden chain of thought looks in a demo.

Further reading

Thinking, Fast and Slow

Thinking, Fast and Slow — Daniel Kahneman’s classic on two modes of thinking — a readable backdrop for why “more deliberation” sometimes helps and sometimes backfires.

A Brief History of Intelligence

A Brief History of Intelligence — Max Bennett’s tour of brain evolution and AI breakthroughs — useful context for today’s bet on machines that spend more compute to reason.

Leave a Comment

Your email address will not be published. Required fields are marked *

Scroll to Top
Aglena