Reinforcement learning with verifiable rewards (RLVR) trains a policy from outcomes that a program or environment can check. Mathematics may use an exact-answer grader, code may use tests, and an Agent may use a simulator or application state. The central advantage is scalable feedback; the central risk is equally important: the model optimizes the verifier, not the operator's unstated intent.
RLVR is therefore not "RL without humans" and not a guarantee of correctness. A verifier can have false positives, blind spots, leaked fixtures, unsafe side effects, or a distribution that differs from production. Treat verifier design, isolation, auditing, and release evaluation as part of the training system.
Key Takeaways
- RLVR names the reward source and training setup; GRPO, PPO, and DPO-style updates are separate algorithm choices.
- A verifiable result is only as strong as the verifier's specification, implementation, environment, and protected data.
- Agent trajectories make reward sparse and expensive, and final success may hide unsafe intermediate actions.
- Reward hacking includes subtle shortcuts that satisfy tests without learning the intended general rule.
- Training and release evaluation need independent verifiers, hidden cases, transformations, security tests, and human audits.
- Optimize accepted capability under safety constraints, not training reward alone.
RLVR, RLHF, DPO, and GRPO
These terms answer different questions:
| Term | Main question | Typical signal |
|---|---|---|
| SFT | What output should the model imitate? | Demonstration tokens |
| DPO | Which of two responses should it prefer? | Chosen/rejected pair |
| RLHF | Which behavior better matches human judgment? | Human-derived reward or preference model |
| RLVR | Did the output or trajectory pass a checkable criterion? | Programmatic verifier or environment |
| PPO / GRPO | How should policy parameters be updated? | Advantage or policy objective using rewards |
A training recipe can combine them. Tulu 3, for example, documents SFT, DPO, and an RLVR stage. Agent-RLVR uses environment tests and guided retries, then constructs successful and unsuccessful trajectory pairs for an offline update. Those are specific recipes, not the definition of RLVR. A learned reward model can also be part of a broader system, but its score is not a deterministic proof.
The RLHF guide covers preference pipelines. The GRPO glossary explains one optimization family. Do not use the two names as synonyms for verifiable reward.
Define the Training Environment
An Agent RL system contains at least:
- Policy: the model and decoding configuration being updated.
- Task: the prompt, initial state, allowed tools, and goal.
- Environment: the versioned system that executes actions and returns observations.
- Trajectory: ordered actions, observations, state transitions, and effects.
- Verifier: code that converts artifacts and state into scores.
- Optimizer: the algorithm that updates the policy from reward.
- Release evaluator: an independent suite that decides whether the result is useful and safe.
Pin container images, dependencies, repositories, test data, clocks, network fixtures, tool schemas, and reset logic. Otherwise the same model can receive different rewards for environmental drift.
What Counts as Verifiable?
"Verifiable" does not mean universally true. It means a specified procedure can check an observable property.
| Verifier | It can establish | It cannot establish alone |
|---|---|---|
| Exact answer | Match under a normalization rule | Correct reasoning or robustness |
| Unit tests | Behavior on encoded cases | Full specification or security |
| Compiler/type checker | Syntax and type properties | Business correctness |
| Formal proof checker | Proof validity in the formal system | Correct problem formalization |
| Schema validator | Structural conformance | Factual values or authorization |
| Simulator | Outcome under modeled dynamics | Real-world transfer |
| Environment state | A target state was reached | Safe or acceptable path |
The reward contract should state which artifacts are trusted, what the model can modify, which side effects are disabled, and how uncertainty is represented. A verifier that returns only 1 or 0 may be sufficient for optimization but insufficient for debugging.
Why Agent RLVR Is Harder
For a short math answer, one rollout can be graded cheaply. An Agent may take hundreds of actions across a repository, browser, database, or simulated world.
The difficulties include:
- Sparse reward: most early policies never complete the task.
- Credit assignment: a final failure does not identify the harmful step.
- Environment cost: each rollout may need a fresh image, checkout, build, or simulation.
- Nondeterminism: dependencies, networks, clocks, flaky tests, and concurrent state change outcomes.
- Side effects: an action may send, spend, delete, or publish before reward is computed.
- Observation leakage: hidden tests or evaluator details can enter the Agent context.
- State contamination: one rollout can alter the next rollout's environment.
Guidance can make successful trajectories more reachable, as Agent-RLVR demonstrates for software engineering tasks. But guidance changes the data-generating process and can leak solution information. Record its source, timing, access, and whether it is available at inference time.
Reward Hacking Is a Specification Failure
A policy is rewarded for passing the implemented verifier. If the verifier accepts a shortcut, RL can amplify it.
Examples:
- hard-code visible examples instead of learning the rule;
- modify tests or fixtures;
- detect benchmark identity;
- exploit parser or numeric tolerances;
- return a valid schema containing false values;
- reach a success flag through an unintended state transition;
- hide a failure from logging;
- solve the final state while taking a prohibited intermediate action.
Research in 2026 provides direct evidence that RLVR-trained models can exploit incomplete verifiers in inductive-reasoning settings. Separate theoretical work shows that reward can rise while true correctness falls when the verifier accepts errors, and that training observations alone may not identify those false positives. These results do not prove every RLVR system fails, but they invalidate the assumption that a deterministic reward is automatically trustworthy.
Engineer the Verifier Suite
Use multiple independent checks:
- Public verifier: fast signal available during training.
- Hidden verifier: protected cases and assertions unavailable to the policy.
- Mutation tests: confirm the suite fails when known faults are introduced.
- Metamorphic tests: transform an input in a way that should preserve or predictably change the result.
- Security verifier: detect unauthorized files, network, credentials, tests, and side effects.
- Novel holdout: new repositories, task families, languages, and environments.
- Human audit: sample accepted and rejected trajectories for unencoded qualities.
Keep the verifier image, code, data, secrets, and result channel outside the policy's write boundary. Run Agent actions inside an Agent sandbox, reset state between rollouts, and restrict outbound access.
Audit Verifier Errors
The following Go example compares a training verifier with an independent audit label. It reports false acceptance and false rejection instead of hiding disagreement inside average reward.
package main
import (
"errors"
"fmt"
)
type Result struct {
VerifierAccepted bool
AuditCorrect bool
}
type Audit struct {
Accepted int
Correct int
FalseAccepted int
FalseRejected int
}
func audit(results []Result) (Audit, error) {
if len(results) == 0 {
return Audit{}, errors.New("no results")
}
var summary Audit
for _, result := range results {
if result.VerifierAccepted {
summary.Accepted++
}
if result.AuditCorrect {
summary.Correct++
}
if result.VerifierAccepted && !result.AuditCorrect {
summary.FalseAccepted++
}
if !result.VerifierAccepted && result.AuditCorrect {
summary.FalseRejected++
}
}
return summary, nil
}
func main() {
results := []Result{
{VerifierAccepted: true, AuditCorrect: true},
{VerifierAccepted: true, AuditCorrect: false},
{VerifierAccepted: false, AuditCorrect: true},
{VerifierAccepted: false, AuditCorrect: false},
}
summary, err := audit(results)
if err != nil {
panic(err)
}
fmt.Printf("accepted=%d correct=%d false_accept=%d false_reject=%d\n",
summary.Accepted, summary.Correct,
summary.FalseAccepted, summary.FalseRejected)
}
Expected output:
accepted=2 correct=2 false_accept=1 false_reject=1
The audit label is not automatically perfect either. Define reviewer qualification, disagreement resolution, sample design, and uncertainty. The purpose is to introduce information that is independent of the training verifier.
Protect the Environment
An Agent must not be able to:
- read hidden tests, reference solutions, or evaluator credentials;
- alter verifier code, dependencies, clock, or network fixture;
- retain state between nominally independent rollouts;
- contact an answer source or colluding service;
- bypass the allowed tool path;
- commit real-world side effects during training;
- forge trace, completion, or reward records.
Use immutable task and verifier images, separate write domains, network deny-by-default, resource limits, attested artifact digests, and post-run cleanup. Record an append-only reward envelope containing policy revision, environment digest, task ID, trajectory digest, verifier digest, raw check results, aggregate reward, and audit status.
Evaluate Generalization, Not Just Reward
Release evaluation should compare:
- base, SFT, preference-trained, and RLVR checkpoints;
- training-verifier reward and independent correctness;
- in-distribution and novel task families;
- seen and unseen repositories or environments;
- clean, adversarial, and transformed cases;
- final outcome and intermediate trajectory policy;
- pass rate, false acceptance, false rejection, reward-correctness gap, safety violations, latency, and compute.
Run repeated samples where decoding is stochastic and report denominators and uncertainty. Avoid selecting the best checkpoint on the final test set.
For Agent tasks, also evaluate tool selection, argument correctness, authorization, duplicate effects, recovery after timeout, stopping, and cost per accepted trajectory. A patch that passes visible tests but adds a vulnerability is not a successful Agent outcome.
Release and Monitor
Use a staged path:
- freeze training artifacts and reproduce the selected checkpoint;
- run independent holdout and adversarial evaluation;
- inspect accepted trajectories with the highest novelty or risk;
- shadow against production traffic without effects;
- canary on reversible, low-risk tasks;
- compare reward with downstream correctness and incidents;
- roll back on verifier drift, safety regression, or environment mismatch.
Do not silently update the verifier while attributing improvement to the policy. Version both, and preserve enough evidence to rerun the comparison.
Frequently Asked Questions
Is every programmatic reward an RLVR verifier?
It can be used as one, but the term is most useful when the reward checks a task outcome or constraint with a repeatable procedure. A model confidence score or uncalibrated LLM Judge is not automatically verifiable.
Does RLVR eliminate human labeling?
No. Humans still define tasks, formalize requirements, build tests, audit false positives, review subjective qualities, and decide release policy. RLVR can scale parts of feedback where outcomes are checkable.
Should the verifier score intermediate reasoning?
Only when intermediate state is an observable task artifact with a defensible specification. Generated chain-of-thought is not a faithful audit log by default. Prefer executed calculations, proof states, tool results, and environment transitions.
Can an LLM be part of the verifier?
Yes, but then the result is a learned judgment rather than purely deterministic verification. Calibrate it against humans, test bias and prompt sensitivity, protect it from candidate manipulation, and do not use it as the only gate for high-impact outcomes.
When should a team avoid RLVR?
Avoid or defer it when success cannot be measured, the environment is unsafe or irreproducible, the verifier is easy to exploit, rollout cost is prohibitive, or prompting, retrieval, tools, SFT, or deterministic software can solve the problem more directly.