TL;DR

A reasoning model is not a single architecture. It is a model or serving system optimized to spend additional inference compute on selected problems. OpenAI disclosed reinforcement learning and test-time scaling for o1, while DeepSeek published a more detailed R1 training recipe. Neither a long answer nor a visible Chain of Thought proves that a model searched correctly. Production systems need matched-budget evaluations, independent verifiers, stopping rules, and a lower-cost fallback path.

Table of Contents

Key Takeaways

  • Separate three layers: foundation architecture, post-training recipe, and inference policy are different engineering decisions.
  • Use only disclosed facts: o1's exact architecture, search procedure, and reward design remain private.
  • Distinguish R1-Zero from R1: the pure-RL experiment is not the complete released R1 recipe.
  • Treat compute as a budget: a longer trace, more samples, and verifier-guided search are different test-time strategies.
  • Measure accepted outcomes: report task success, p95 latency, total cost, abstention, and verifier errors under a fixed protocol.
  • Do not audit from prose: visible reasoning is generated text, not a complete execution or safety log.

What Reasoning Models Are

A reasoning model is best understood as an operational category: it is trained or served to allocate more computation to tasks that benefit from multi-step work. The term does not identify a unique Transformer variant, a mandatory hidden scratchpad, or a guaranteed search algorithm.

The System 1 and System 2 analogy is useful for describing fast versus deliberate behavior, but it is not a mechanistic account of a neural network. A standard large language model can solve difficult problems, and a reasoning model can confidently fail.

The three layers that should not be conflated

Layer Question it answers Publicly observable evidence Common overclaim
Foundation architecture What network produces token probabilities? Model report, weights, configuration "Long reasoning proves a new architecture"
Post-training recipe Which behaviors were reinforced after pretraining? SFT/RL stages, reward design, data description "The model learned entirely without examples"
Inference policy How is compute allocated for this request? Effort controls, samples, token usage, tool calls "One long response is a tree search"

This distinction matters when comparing OpenAI o1 with DeepSeek R1. OpenAI published behavioral and evaluation evidence but not a complete architecture or training recipe. DeepSeek published weights for several models and a technical report with substantially more training detail. The difference in disclosure does not, by itself, prove which system is better for a workload.

flowchart LR A["Foundation model"] --> B["Post-training policy"] B --> C["Inference strategy"] C --> D["Candidate answer"] D --> E["Independent verification"] E --> F["Accept, retry, or abstain"]

What OpenAI Disclosed About o1

OpenAI's public o1 report supports three claims: o1 was trained with reinforcement learning, it used additional internal computation before answering, and its measured performance increased with more training-time and test-time compute under the reported evaluations. It does not disclose enough information for an exact "o1 architecture" reconstruction.

The report also illustrates why decoding policy must be part of every benchmark claim. On its 2024 AIME evaluation, OpenAI reported 74% with one sample, 83% with majority voting over 64 samples, and 93% when a learned scoring function reranked 1,000 samples. Those are three different systems and budgets, not three measurements of the same request path.

Supported by the public report Not established by the public report
Reinforcement learning was used to improve reasoning behavior The exact base-model architecture or parameter count
More test-time compute improved selected reported benchmarks A universal monotonic inference scaling law
Hidden reasoning was used before the final answer A faithful, complete, externally auditable trace
Single-sample, consensus, and reranked results differed A specific PRM, MCTS, beam search, or verifier implementation

This is the practical rule: describe o1 by what OpenAI measured and disclosed, not by reverse-engineering fluent output. Current API controls and token accounting are versioned product contracts; check the current OpenAI reasoning guide before implementing a budget or parser.

How DeepSeek R1 Was Trained

DeepSeek R1 is easier to study because its technical report separates the experimental R1-Zero path from the full R1 pipeline. Collapsing both into "R1 was trained with pure RL" removes the most important engineering lesson.

R1-Zero tested rule-based reinforcement learning

DeepSeek-R1-Zero started from DeepSeek-V3-Base and skipped a supervised fine-tuning warm start before reinforcement learning. The report describes GRPO, which estimates relative advantages within a group of sampled outputs rather than training a separate critic in the PPO style.

For the reported mathematical, coding, and logical tasks, R1-Zero used two rule-based reward components:

  1. Accuracy reward checked final answers with deterministic rules or code tests.
  2. Format reward checked whether reasoning and answers used the required delimiters.

The authors explicitly say they did not use neural outcome or process reward models for this R1-Zero reasoning stage because of reward-hacking risk and retraining complexity. This fact is easy to lose when generic explanations insert a process reward model into the R1 architecture.

R1-Zero also exposed the limits of the experiment. The report describes readability problems, language mixing, and weak general-purpose behavior. Longer responses and self-reflective phrases were observed behaviors; they were not proof that every intermediate claim was correct.

The released R1 used a multi-stage recipe

The full DeepSeek-R1 pipeline added data and alignment stages rather than shipping the R1-Zero recipe unchanged:

  1. Cold-start supervised fine-tuning established readable reasoning patterns.
  2. Reasoning-oriented reinforcement learning improved performance on verifiable tasks.
  3. Rejection sampling and supervised fine-tuning incorporated successful reasoning trajectories plus non-reasoning data.
  4. A second reinforcement-learning stage targeted helpfulness, safety, and broader behavior.
  5. Distillation transferred generated reasoning data into smaller dense models.
Artifact Training role What it proves What it does not prove
R1-Zero Study reasoning behavior from rule-reward RL without an SFT warm start RL can substantially improve the reported verifiable tasks Pure RL is sufficient for a polished general assistant
DeepSeek-R1 Multi-stage reasoning and alignment model Cold start, RL, rejection sampling, SFT, and alignment can be combined Every hosted R1-like service uses the identical pipeline
Distilled models Smaller models trained on generated reasoning data Distillation can transfer useful behavior Distillation reproduces the teacher's internals

The same report lists process reward models and Monte Carlo tree search among approaches that did not deliver the desired result in the authors' setup. PRMs and MCTS are valid research techniques, but they should not be presented as confirmed internal components of DeepSeek-R1. For the base network context, see Mixture of Experts Architecture Explained; for preference optimization context, see RLHF and Preference Learning.

How Test-Time Compute Changes Inference

Test-time compute changes how much work a system performs after receiving a prompt, but "thinking longer" can refer to several distinct mechanisms. Comparing them only by output-token count hides the actual intervention.

Strategy Where extra compute goes Selection signal Primary risk
Longer single trajectory More sequential tokens in one sample Model's own continuation policy Error accumulation and circular reasoning
Parallel sampling Independent complete candidates Majority or task-specific aggregation Correlated candidates create false confidence
Best-of-N Multiple candidates Outcome verifier or reward model Verifier selects polished but wrong answers
Verifier-guided search Partial candidates and branching Process or state scorer Search exploits scorer weaknesses
Tool-assisted loop Calls to tests, retrieval, or solvers External tool results Tool errors, permissions, and stale state

These mechanisms produce different quality, latency, and cost curves. A model with a longer private trace is not equivalent to a service that samples 64 answers, and majority voting is not equivalent to selecting with a learned verifier.

The test-time scaling study by Snell and colleagues found that the best allocation depended on both question difficulty and the base model. Their compute-optimal strategy outperformed a Best-of-N baseline with substantially less compute in the specific PaLM 2 and MATH setup, while the hardest questions sometimes benefited little from additional test-time work. This is an experimental result with explicit conditions, not a universal law.

For implementation patterns such as self-consistency, search, and adaptive budgets, use the separate Test-Time Compute Engineering Guide. The rest of this article focuses on evaluating a reasoning product rather than rebuilding search algorithms.

How to Evaluate Reasoning Models

A useful reasoning-model evaluation holds the task, model snapshot, tools, and acceptance rule constant while varying the compute policy. Otherwise, a higher score may come from a larger model, more samples, privileged tools, or a different grader rather than better reasoning.

Freeze the evaluation contract

Record these fields for every run:

  • dataset version and contamination controls;
  • model and provider snapshot;
  • system prompt and user prompt template;
  • reasoning effort, temperature, maximum output, and sample count;
  • tool definitions, permissions, timeouts, and fixture versions;
  • answer extractor, verifier version, and abstention rule;
  • retries, cache policy, concurrency, region, and timestamp.

Compare matched operating points

At minimum, compare four paths on the same held-out tasks:

  1. a lower-cost model or effort setting with one sample;
  2. the reasoning model with one bounded response;
  3. multiple candidates aggregated by consensus;
  4. multiple candidates selected by an independent verifier.

Do not compare a single-sample baseline with a 64-sample reasoning result and attribute the full difference to model quality. Report both model identity and inference policy.

Use metrics that preserve meaning

Metric What it measures Frequent misuse
Pass@1 Success of the delivered first candidate Reporting it after hidden retries
Pass@k Probability that at least one of k candidates succeeds Treating it as user-visible accuracy without a selector
Consensus accuracy Accuracy of an aggregation rule Assuming agreement implies independence
Verifier-selected accuracy Quality after a stated selector Ignoring verifier false positives
Cost per accepted task Total model, tool, retry, and review cost divided by accepted successes Reporting token price alone
p95 and p99 latency Tail response time for completed and failed requests Reporting only mean latency
Abstention quality Whether rejected cases are actually riskier Counting every refusal as safety

Slice results by task family, difficulty, input length, language, tool availability, and failure consequence. A routing policy needs calibrated slice-level evidence, not one blended benchmark average.

A Verifier Contract in Go

A production verifier should evaluate a narrow, testable output contract and be independent from the model's prose. The following Go program accepts only an exact integer in a FINAL: field, rejects every invalid candidate, then chooses the lowest-cost verified candidate with latency as a tie-breaker.

go
package main

import (
	"errors"
	"fmt"
	"sort"
	"strconv"
	"strings"
)

type Candidate struct {
	ID         string
	Output     string
	CostMicros int
	LatencyMS  int
}

type VerifiedCandidate struct {
	Candidate
	Answer int
}

func verifyInteger(output string, expected int) (int, error) {
	const prefix = "FINAL:"

	if !strings.HasPrefix(output, prefix) {
		return 0, errors.New("missing FINAL field")
	}

	answer, err := strconv.Atoi(strings.TrimSpace(strings.TrimPrefix(output, prefix)))
	if err != nil {
		return 0, fmt.Errorf("invalid integer: %w", err)
	}
	if answer != expected {
		return 0, fmt.Errorf("expected %d, got %d", expected, answer)
	}

	return answer, nil
}

func selectVerified(candidates []Candidate, expected int) (VerifiedCandidate, error) {
	verified := make([]VerifiedCandidate, 0, len(candidates))

	for _, candidate := range candidates {
		answer, err := verifyInteger(candidate.Output, expected)
		if err != nil {
			continue
		}
		verified = append(verified, VerifiedCandidate{
			Candidate: candidate,
			Answer:    answer,
		})
	}

	if len(verified) == 0 {
		return VerifiedCandidate{}, errors.New("no candidate passed verification")
	}

	sort.Slice(verified, func(i, j int) bool {
		if verified[i].CostMicros == verified[j].CostMicros {
			return verified[i].LatencyMS < verified[j].LatencyMS
		}
		return verified[i].CostMicros < verified[j].CostMicros
	})

	return verified[0], nil
}

func main() {
	candidates := []Candidate{
		{ID: "candidate-a", Output: "FINAL: 41", CostMicros: 2100, LatencyMS: 420},
		{ID: "candidate-b", Output: "FINAL: 42", CostMicros: 3200, LatencyMS: 700},
		{ID: "candidate-c", Output: "FINAL: 42", CostMicros: 4200, LatencyMS: 610},
	}

	selected, err := selectVerified(candidates, 42)
	if err != nil {
		fmt.Println("abstain:", err)
		return
	}

	fmt.Printf(
		"selected=%s answer=%d cost_micros=%d latency_ms=%d\n",
		selected.ID,
		selected.Answer,
		selected.CostMicros,
		selected.LatencyMS,
	)
}

Expected output:

text
selected=candidate-b answer=42 cost_micros=3200 latency_ms=700

This is an inference-time acceptance example, not a reconstruction of the R1 reward function. Real verifiers should run unit tests, schema validation, constraint solvers, policy checks, or human review depending on the task. If correctness cannot be checked reliably, generating more candidates may only create more persuasive errors.

Failure Modes and Production Controls

Reasoning systems fail when extra compute amplifies a weak proposal, a biased verifier, or an unsafe action loop. Each failure mode needs an observable control outside the model's narrative.

Failure mode Observable symptom Required control
Overthinking Higher effort increases latency but lowers task success Per-slice effort cap and lower-effort fallback
Correlated sampling Many candidates repeat the same wrong answer Prompt/model diversity and executable checks
Verifier gaming Selected answer scores well but fails real acceptance tests Hidden holdout tests and verifier red-team cases
Budget exhaustion Request ends before a usable answer Reserve output budget, stop early, and return explicit failure
Tool-loop drift Repeated or unauthorized actions Idempotency keys, least privilege, call limits, and audit events
Reasoning leakage Sensitive prompt or data appears in stored traces Data minimization, redaction, and retention policy
Unfaithful rationale Explanation omits an influential hint or tool result Audit inputs, outputs, tool calls, and evidence instead of prose

Anthropic's controlled experiments found that visible reasoning often failed to mention hints that influenced answers, including in experiments involving reward hacking. The study used artificial hinting tasks and selected models, so it does not prove that every trace is unfaithful. It does show why Chain of Thought cannot serve as a complete security audit log.

A minimum release gate

A reasoning path should ship only when all of these statements are true:

  • it beats the lower-cost baseline on a versioned, representative holdout set;
  • the gain survives matched-budget and matched-tool comparisons;
  • p95 latency, timeout rate, and cost per accepted task meet the service objective;
  • verifier false positives and false negatives are measured on adversarial cases;
  • no-valid-answer behavior is explicit and tested;
  • tool permissions, retries, and side effects are bounded;
  • monitoring can attribute regressions to model, prompt, verifier, or routing changes.

Canary the new policy before broad rollout. Roll back on accepted-answer quality, not on a model-reported confidence score.

How to Choose a Reasoning Model

Choose a reasoning model from the workload contract, not from a benchmark headline. OpenAI o1 and DeepSeek R1 represent different disclosure and deployment options, while a standard model may still be the best operating point for routine work.

Requirement Prefer testing first Why
Managed API and minimal model operations A provider reasoning API Provider owns serving, updates, and capacity
Weight access, private deployment, or runtime control An appropriate open-weight R1-family checkpoint You can inspect artifacts and control serving
Tight latency or high-volume routine requests Lower-effort or standard model Extra reasoning cost may not improve acceptance
Verifiable math, code, or constrained planning Reasoning model plus external verifier Objective checks make additional candidates useful
High-stakes open-ended advice Retrieval, tools, and expert review around any model Reasoning text does not establish factual reliability

Run a routing experiment instead of creating a permanent task blacklist. A practical policy starts with the cheaper path, escalates only on measured high-value slices, and abstains when the verifier cannot establish an acceptable answer. For provider-specific switches, see Hybrid Reasoning Models in Practice.

Frequently Asked Questions

Is a reasoning model a new neural architecture

Not necessarily. "Reasoning model" usually describes a trained behavior or inference product. Public evidence may reveal the base architecture, post-training method, or serving controls to different degrees. Unless a provider publishes the network details, calling the system a new architecture is speculation.

Does DeepSeek R1 use pure reinforcement learning

R1-Zero is the pure-RL experiment in the report: it starts from a base model without a supervised reasoning warm start. The full DeepSeek-R1 recipe adds cold-start data, supervised fine-tuning, rejection sampling, another reinforcement-learning stage, and distillation for smaller models.

Is more reasoning effort always better

No. Extra compute is useful only when the model can generate better candidates and the system can identify them. Easy tasks may only become slower, while very hard tasks can remain beyond the proposal model or verifier. Measure the complete quality-cost-latency curve.

Is majority voting the same as verification

No. Majority voting measures agreement among samples. An independent verifier checks a task property such as passing tests or satisfying constraints. If samples share the same systematic error, consensus can be confidently wrong.

Can visible Chain of Thought be logged for compliance

It can be supplementary diagnostic data if policy permits, but it is not a sufficient compliance record. Log the request, model and prompt versions, tool calls, retrieved evidence, policy decisions, final output, verifier result, and human approvals. Minimize or redact reasoning text that may contain sensitive data.

Summary

Reasoning models make inference compute a controllable engineering resource, not a guarantee of truth. The defensible comparison is not "which model thinks longer," but which fixed model-and-policy combination produces more independently verified outcomes within your latency, cost, and safety limits.

Sources and Further Reading