TL;DR

Test-time compute (TTC) is the computation spent while answering a request. Engineers can allocate more of it to longer reasoning, alternative candidates, verification, and search. Additional work helps only when it produces useful candidates and the system can select them. Budget generation and verification together, measure selected-answer accuracy, and make exhaustion an explicit outcome.

Contents

What test-time compute changes

TTC changes how much work a system performs at inference and how it spends that work. It does not require updating model weights. A longer reasoning trajectory, five independently sampled answers, and a search tree with a verifier are different allocations of inference compute.

Ordinary autoregressive LLM generation already consists of repeated decoding steps. The prompt is processed, then successive tokens are generated, usually with a KV cache. Calling this a “single forward pass” hides the main source of variable generation cost.

Distinguish three layers:

Layer What changes Example
Model training Parameters and learned reasoning behavior Training a model to solve multi-step problems
Inference policy Work performed for one request Longer generation or sampling multiple solutions
Application controller Which work to start, accept, or stop Running tests and revising a rejected candidate

Chain-of-Thought prompting can elicit intermediate reasoning, but it is not a guarantee of correctness. A reasoning API can expose an effort setting without exposing its internal search implementation. Neither a long explanation nor a visible “let me reconsider” sentence proves that verification happened.

Snell et al.'s test-time scaling study investigates adaptive proposal strategies and verifier-guided search. Its central engineering lesson is that the useful allocation depends on problem difficulty. Its reported compute advantages belong to the studied models, tasks, and budgets; they are not a universal multiplier for production workloads.

Choose a strategy by its decision rule

Choose the strategy according to the evidence available for selecting an answer. Greater algorithmic complexity does not imply greater accuracy.

Strategy What receives extra compute How an answer is chosen Main failure mode
Longer trajectory One generation Model produces a final answer Longer output repeats the same mistake
Self-consistency Multiple complete reasoning paths Aggregate equivalent final answers Shared errors win the vote
Best-of-N Multiple complete candidates plus scoring Highest acceptable verifier score Scorer favors a wrong candidate
Sequential revision Feedback and edits to a candidate Recheck each revision A repair breaks working behavior
Tree-of-Thought Branches of intermediate proposals Search with evaluation and backtracking Wrong pruning removes useful paths
MCTS Selection, expansion, evaluation, and backups A defined terminal selection policy Noisy value estimates attract more search

The original Self-Consistency paper samples reasoning paths and aggregates their answers. These samples share a model and prompt, so their errors need not be independent. Answer normalization also matters: 42, 42.0, and “42 meters” may or may not mean the same thing under the task's contract.

Best-of-N scores complete candidates. Beam search repeatedly expands partial sequences or states and retains a bounded set. Generating five complete answers and selecting one is not, by itself, beam search.

Tree-of-Thought explicitly explores intermediate thoughts. For a search state to be useful, define what it contains, which transitions are legal, and what makes it terminal. A node that merely contains a persuasive paragraph is difficult to validate.

For example, a scheduling task can represent a state as assigned jobs, remaining jobs, and capacity constraints. A verifier can reject an impossible assignment before spending more on that branch. In an essay-writing task, such a crisp state and verifier may not exist; measured rubric-based revision may be more appropriate.

flowchart TD A["Request and acceptance contract"] --> B["Reserve work budget"] B --> C["Generate or expand candidates"] C --> D["Parse and verify"] D --> E{"Evidence sufficient"} E -->|Yes| F["Return selected result and evidence"] E -->|No| G{"Budget remains"} G -->|Yes| B G -->|No| H["Return unresolved or escalate"]

Account for the complete request

A TTC budget must cover generation, selection, verification, and retries. Counting only the final response systematically understates cost.

For token-priced calls, a simplified accounting identity is:

total cost = sum(input usage × applicable input rate + output usage × applicable output rate) + tool costs

Apply cached-input or other pricing categories according to the provider's actual contract. The formula is an accounting structure, not a price quote. Log the model and pricing version used for reconciliation.

The OpenAI reasoning guide specifies that reasoning tokens are included in output_tokens. They are a detail within that total, not another quantity to add to it. Do not equate output_tokens - reasoning_tokens with an exact count of visible text either: the API can include other nonvisible output formatting tokens.

max_output_tokens caps reasoning and other output tokens together. A response can exhaust that cap before producing a visible answer. An effort setting controls reasoning behavior for supported models; it is not a hard dollar limit. Parameter support and usage fields must be checked for the chosen model and endpoint.

Use several limits because they solve different problems:

Limit What it bounds What it does not bound
Maximum attempts Number of initiated calls Tokens in each call
Per-call output cap Output token allocation Input cost or tool cost
Total reserved output Concurrent or sequential output commitments Monetary cost across different rates
Deadline How long the client waits Guaranteed cancellation of server-side work
Search depth and node count Tree expansion Cost of evaluating each node
Tool permissions and quotas External effects and tool usage Model answer quality

Reserve a call's upper bound before dispatch. Reconcile it with trusted usage afterward. If a timeout leaves usage unknown, keep the reservation until reconciliation; treating missing usage as zero can let retries exceed the intended budget. Parallel workers need an atomic shared reservation, not a separate budget copy each.

Verifier calls compete with candidate generation. The Art of Scaling Test-Time Compute studies that tradeoff under a fixed total budget and finds settings where self-consistency is more efficient than a generative reward model. This supports measuring verifier cost, not discarding verification in every application.

A bounded self-consistency controller in Go

The following Go 1.23 example demonstrates output reservations, final-answer parsing, explicit failure, and a stopping rule. It uses scripted responses so the controller can be reproduced without API credentials. It does not run an LLM, estimate accuracy, or impose a total monetary budget.

The task contract accepts a signed base-10 integer that fits in int64. Truncated responses do not vote. A selection needs a unique leading answer and a minimum number of votes. Early stopping occurs only when the leader cannot be overtaken or tied even if every remaining attempt votes for the runner-up.

Save the code as main.go and run go run main.go.

go
package main

import (
	"context"
	"fmt"
	"strconv"
	"strings"
	"time"
)

type Sample struct {
	Answer     string
	Output     int
	UsageKnown bool
	Complete   bool
}

type Generate func(context.Context, int) (Sample, error)

type Config struct {
	MaxAttempts int
	PerCall     int
	TotalOutput int
	MinVotes    int
}

type Result struct {
	Answer   string
	Selected bool
	Attempts int
	Valid    int
	Charged  int
	Stop     string
}

func leader(votes map[string]int) (string, int, int) {
	answer, first, second := "", 0, 0
	for value, count := range votes {
		if count > first {
			answer, second, first = value, first, count
		} else if count > second {
			second = count
		}
	}
	return answer, first, second
}

func solve(ctx context.Context, cfg Config, generate Generate) (Result, error) {
	var result Result
	if cfg.MaxAttempts < 1 || cfg.PerCall < 1 || cfg.TotalOutput < 1 ||
		cfg.MinVotes < 1 || cfg.MinVotes > cfg.MaxAttempts || generate == nil {
		return result, fmt.Errorf("invalid configuration")
	}
	votes := make(map[string]int)
	for result.Attempts < cfg.MaxAttempts {
		if err := ctx.Err(); err != nil {
			return result, err
		}
		remaining := cfg.TotalOutput - result.Charged
		if remaining == 0 {
			result.Stop = "output_budget"
			break
		}
		cap := min(cfg.PerCall, remaining)
		result.Charged += cap
		result.Attempts++
		sample, err := generate(ctx, cap)
		if sample.UsageKnown {
			if sample.Output < 0 || sample.Output > cap {
				return result, fmt.Errorf("usage outside reserved cap")
			}
			result.Charged -= cap - sample.Output
		}
		if err != nil {
			return result, fmt.Errorf("generation failed: %w", err)
		}
		if err := ctx.Err(); err != nil {
			return result, err
		}
		if !sample.Complete {
			continue
		}
		value, err := strconv.ParseInt(strings.TrimSpace(sample.Answer), 10, 64)
		if err != nil {
			continue
		}
		votes[strconv.FormatInt(value, 10)]++
		result.Valid++
		_, first, second := leader(votes)
		if first >= cfg.MinVotes &&
			first > second+cfg.MaxAttempts-result.Attempts {
			result.Stop = "locked_plurality"
			break
		}
	}
	if result.Stop == "" {
		result.Stop = "attempt_limit"
	}
	answer, first, second := leader(votes)
	if first >= cfg.MinVotes && first > second {
		result.Answer, result.Selected = answer, true
	}
	return result, nil
}

func main() {
	ctx, cancel := context.WithTimeout(context.Background(), time.Second)
	defer cancel()
	replies := []string{"42", "41", "42", "42", "41"}
	index := 0
	generate := func(ctx context.Context, cap int) (Sample, error) {
		if err := ctx.Err(); err != nil {
			return Sample{}, err
		}
		if cap < 12 || index >= len(replies) {
			return Sample{}, fmt.Errorf("scripted response unavailable")
		}
		reply := replies[index]
		index++
		return Sample{reply, 12, true, true}, nil
	}
	result, err := solve(ctx, Config{5, 32, 160, 3}, generate)
	if err != nil {
		fmt.Println("error:", err)
		return
	}
	fmt.Printf("answer=%s selected=%t attempts=%d valid=%d charged=%d stop=%s\n",
		result.Answer, result.Selected, result.Attempts, result.Valid,
		result.Charged, result.Stop)
}

Expected output:

text
answer=42 selected=true attempts=4 valid=4 charged=48 stop=locked_plurality

After four calls the leading answer has three votes, the other has one, and only one attempt remains. The leader cannot be tied. selected=true means the declared selection rule passed; it does not mean an external oracle verified 42.

Important behaviors to test before adapting this controller:

  • 42, 41 with two attempts and one required vote ends unresolved because of a tie.
  • Malformed and incomplete responses consume budget but do not vote.
  • A known output count refunds unused reservation; unknown usage retains the entire cap.
  • An API error stops this example and returns no selected answer. A retry policy must reserve a new attempt.
  • Budget exhaustion can still yield a unique plurality if the minimum vote rule passed; the stop reason records the exhausted resource. Require external verification if plurality is insufficient for the application.

The live adapter must pass the output cap to the provider, propagate context cancellation, inspect completion status, and report authoritative usage. Input limits, tool costs, admission control, rate limits, and persistent billing reconciliation remain separate requirements.

Make search and verification trustworthy

A verifier is valuable when it supplies evidence that the generator does not reliably supply on its own. Its scope must be explicit.

Evidence What a pass establishes Remaining uncertainty
Parser or schema Output satisfies a structural contract Values can be false
Arithmetic or constraint check Encoded conditions hold Conditions may be incomplete
Test suite Tested behavior matches expectations Untested behavior and test quality
Retrieved source A source contains relevant evidence Freshness, authority, and entailment
Model judge A learned evaluator prefers the answer Bias, calibration, shared blind spots

Outcome verification evaluates a completed candidate. Process verification evaluates intermediate steps and can guide pruning. Neither is automatically superior: an early false rejection can delete the only useful branch, while a weak outcome scorer can select an attractive but incorrect answer.

For tree search, enforce depth during expansion, cap total nodes, and define terminal validity separately from heuristic value. MCTS visit counts describe search allocation under a particular policy; they are not calibrated probabilities of truth. Returning the most visited unfinished paragraph is not equivalent to returning a validated solution.

Keep generated code away from credentials and host privileges. A subprocess with a timeout is not a security sandbox. Code verification needs an isolated execution environment with resource, filesystem, and network controls, plus tests that generated code cannot quietly rewrite.

When applying feedback to an existing answer, preserve checked candidates and recheck each revision. The self-correction engineering guide covers acceptance, regression, and stopping in that loop.

Evaluate adaptive budgets

An adaptive policy should demonstrate higher selected-answer quality under a stated cost and latency envelope. “The model said this was hard” is a routing feature to evaluate, not ground truth.

Begin with a fixed one-call baseline, then compare a longer trajectory, self-consistency, verifier-based selection, and revision. Keep the model versions, task set, prompts, answer normalization, and grading policy fixed wherever the comparison requires them.

Measure:

Measurement Why it matters
Selected-answer accuracy What the user actually receives
Candidate-set success Whether any sampled candidate was correct
Selection gap Correct candidates exist but the selector misses them
Coverage and accuracy among answered requests Whether abstention merely hides difficult tasks
Cost per request and per correct answer Includes failed requests and verification
P50/P95 latency and deadline failures Exposes queueing and slow branches
Performance by task difficulty and domain Prevents an aggregate score from hiding regressions

Candidate-set success is an oracle view: it assumes access to correctness labels. It should not be presented as deployable accuracy. Record both coverage and unconditional success so a policy that answers only easy questions cannot appear universally better.

Tune thresholds on development data and evaluate on a separate held-out set. If the same benchmarks repeatedly determine prompts, stopping rules, and verifier thresholds, the benchmark itself becomes part of development.

A useful diagnostic is the selection gap. If more samples contain correct answers but deployed accuracy stays flat, spending on a better selector may help. If candidates repeat the same unsupported premise, better input evidence or a more capable model may matter more than additional sampling. The reasoning-model comparison discusses training and evaluation boundaries; context engineering addresses evidence supplied to the model.

FAQ

Is ordinary LLM inference a single forward pass?

No. Autoregressive generation repeatedly runs the model to produce successive tokens, usually reusing a KV cache. Test-time compute changes the inference policy or budget through longer trajectories, multiple candidates, verification, or search.

Does majority voting provide a probability of correctness?

No. Vote share measures agreement among sampled answers. Correlated errors, parsing rules, and task ambiguity can produce a confident-looking majority that is wrong. Calibrate the complete selection policy on held-out tasks.

How should reasoning tokens be counted?

Follow the provider's usage contract. In the OpenAI Responses API, reasoning tokens are included in output_tokens, so adding them again double-counts output usage. Account separately for input, output, verifier calls, and tools.

When should test-time search stop?

Stop when a validated acceptance rule passes or when the request, token, time, depth, or node budget is exhausted. Return an explicit unresolved result when evidence is insufficient; a high search score alone does not establish correctness.

Can extra inference compute replace a larger model?

Sometimes, under a particular task distribution and budget. The result depends on candidate quality, problem difficulty, and selection accuracy. Compare complete systems at equal cost and latency constraints rather than assuming a universal replacement.

Sources and further reading