SYNTHESIS NOTE
Topics›Reinforcement Learning›this note

Can adversarial critics replace task-specific verifiers for reasoning?

Explores whether an adversarial game between policy and critic can substitute for explicit verifiers in RL-based reasoning training. Matters because many domains lack the task-specific validators that make current reasoning RL possible.

Synthesis note · 2026-02-22 · sourced from Reinforcement Learning

A fundamental limitation of RL for reasoning: RLVR requires task-specific verifiers (math checkers, code test suites) that don't exist for many reasoning-intensive domains. Expert demonstrations are abundant (Stack Exchange answers, domain expert explanations) but SFT on demonstrations doesn't produce the reasoning behaviors that large-scale RL training elicits. RARO bridges this gap using inverse reinforcement learning.

The mechanism is an adversarial game. A policy learns to produce expert-like answers via explicit CoT reasoning. A relativistic critic learns to discriminate between expert and policy answers via pairwise comparison. Both are trained jointly and continuously via RL, requiring careful stabilization techniques. The critic's discrimination signal serves as the reward for the policy — when the critic can't distinguish policy from expert, the policy has learned expert-level reasoning.

The results are significant: RARO outperforms strong verifier-free baselines on Countdown, DeepMath, and Poetry Writing, and enjoys the same robust scaling trends as RL with verifiers. This means the scaling properties of RLVR are not specific to verifiable rewards — they emerge from the RL training dynamics themselves, with the adversarial critic providing a sufficient substitute for ground-truth verification.

This extends the frontier of RL-for-reasoning to any domain with expert demonstrations. Since Does critiquing errors teach deeper understanding than imitating correct answers?, RARO leverages a similar mechanism — the adversarial training forces the model to develop genuine reasoning rather than surface-level imitation, because the critic can distinguish superficial pattern matching from actual expert-like problem solving.

VeriFree as a second verifier-free approach: VeriFree takes a different route to the same goal — extending R1-Zero-style RL training to domains without rule-based verifiers. Instead of an adversarial critic, VeriFree generates only the reasoning trace and concatenates it with the reference answer, then evaluates the likelihood of the reference answer conditioned on both. This likelihood serves as both a reward signal for policy gradients on the reasoning trace and a weighting term for supervised training. VeriFree is architecturally simpler than RARO (no adversarial game) and eliminates the need for even a model-based verifier, reducing compute overhead. See Can reasoning improvement work without answer verification?. The two approaches bracket the design space: RARO uses adversarial dynamics for richer signal, VeriFree uses reference-conditioned likelihood for simplicity.

Inquiring lines that read this note 40

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Why does polished AI output gain credibility despite fundamental verifiability problems? How do educators verify student capability when AI can produce indistinguishable work? Which reinforcement learning modifications most improve dialogue quality in language models? How does policy entropy collapse limit scaling of reasoning-focused reinforcement learning? Can AI systems evade safety evaluations through reasoning manipulation? How can evaluations be made robust against model reward hacking? How can we reduce inherent biases in LLM-based evaluation judges? Can external verification systems adequately replace learned reasoning in AI outputs? Why do multi-agent systems reach premature consensus without genuine deliberation? How does RLHF training shape models to prioritize agreement over accuracy? Can AI systems achieve real improvement without external human feedback? Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? Why does AI verification capability persistently exceed generation capability? How do models learn from self-generated outputs without cascading failures? Can humans reliably detect and resist AI-generated misinformation? Can minimal training unlock latent reasoning already present in base models? Should governance of agentic AI systems be runtime or design-time? How do reward models systematically fail to represent diverse human preferences? How reliably can humans and AI detectors identify machine-generated text? How do evaluation environment design choices affect AI security? How does optimization for reward create emergent misalignment in language models?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
18 direct connections · 108 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

inverse rl from demonstrations enables reasoning training without task-specific verifiers