INQUIRING LINE

If an AI can write its own 'rules of success' for training another AI, do humans still need to design them by hand?

Can LLMs solve automated reward design without task-specific prompting or templates?

This explores whether a language model can write good reward functions (the scoring rules that tell a reinforcement learning agent what counts as success) on its own, without a human crafting special prompts or fill-in-the-blank templates for each new task.


This explores whether LLMs can take over the job of writing reward functions for reinforcement learning in general, rather than needing a human to hand-tailor the setup for each task. The corpus's strongest answer is a qualified yes. The headline case is EUREKA, where GPT-4 writes reward code, the code is used to train agents, and the training statistics are fed back so the model can revise its rewards over several rounds of evolutionary search. Those evolved rewards beat human-designed ones on 83% of 29 RL tasks, with a 52% average improvement, and they produced dexterous manipulation skills nobody had hand-specified Can language models design better reward functions than humans?. The method is general: the same loop runs across many tasks without per-task prompt engineering.

The interesting detail is how the generality is achieved. The LLM doesn't get it right in one shot. What makes the approach work without templates is the feedback loop: the model writes a reward, watches what it produces, reflects, and rewrites. MEDIC takes a different route to the same goal. The LLM first solves a simplified, deterministic version of the problem, then turns that plan into shaping rewards for the real, messier task, and a model-based critic checks its output before it's used Can LLMs design reward functions for reinforcement learning?. Both systems replace hand-crafted prompts with structure: an outer loop of search or verification that catches the LLM's mistakes. So the more accurate claim is "no task-specific templates, but a general-purpose scaffolding of feedback."

A lot of the corpus has quietly moved the question somewhere else. Instead of having the LLM *write* the reward, many approaches have the LLM *be* the reward. Reward models that reason step by step before scoring get better as they're given more thinking time Can reward models benefit from reasoning before scoring?. Breaking a vague goal into a checklist of sub-criteria makes subjective tasks gradable Can breaking down instructions into checklists improve AI reward signals?. Giving an LLM judge reference answers makes it about as reliable as a trained reward model Can reference examples make LLM judges reliable enough for self-improvement?. Other work skips designed rewards entirely: it pulls reward signals out of tree search outcomes Can tree search replace human feedback in LLM training?, out of information theory Can we reward reasoning steps without human annotation?, or out of existing metrics like recommendation scores used as a black box Can recommendation metrics train language models directly?. A further twist is that a single number may simply be too thin. Models stuck on a performance plateau improve when they get written critiques explaining *why* they failed Can natural language feedback overcome numerical reward plateaus?.

Two cautions push back on the optimistic picture. First, a reward that scores well can still miss what you meant. Socher argues that reward hacking persists because AI optimizes the literal specification rather than the intent, and an automated reward designer is just another place where that gap can open Why do AIs keep gaming rewards instead of serving intent?. Second, and more surprising, the reward itself may matter less than people assume. For LLM reasoning training, rewards that are spurious or even random work nearly as well as correct ones, because RL mostly brings out skills the model already learned in pretraining What does reward learning actually do to model reasoning?. If that holds more broadly, part of the case for "good automated reward design" rests on the base model, not the reward.

The corpus is thin on direct head-to-head tests of template-free versus template-based reward generation, so how far this generalizes outside robotics simulation is still open. The takeaway is this: LLMs can design rewards without hand-holding, but only when they're wrapped in a loop that lets them see the consequences of what they wrote.


Sources 11 notes

Can language models design better reward functions than humans?

EUREKA uses evolutionary search and training-statistic feedback to evolve reward functions written by GPT-4, outperforming human-designed rewards on 83% of 29 RL benchmarks with 52% average improvement, including novel dexterous manipulation skills.

Can LLMs design reward functions for reinforcement learning?

MEDIC shows that LLMs can generate effective reward shaping functions by first solving a deterministic, simplified version of the RL problem, then converting the resulting plan into shaping rewards for the original stochastic task. A model-based critic validates LLM outputs before deployment.

Can reward models benefit from reasoning before scoring?

Three independent teams (RRM, RM-R1, DeepSeek-GRM) discovered that adding chain-of-thought reasoning before reward scoring enables adaptive test-time compute scaling for evaluation. Reasoning-based approaches raise the capability ceiling of reward models beyond what outcome-based evaluation achieves.

Can breaking down instructions into checklists improve AI reward signals?

RLCF and RaR methods decompose instruction quality into verifiable sub-criteria, improving performance on benchmarks like FollowBench and HealthBench. This decomposition principle reduces overfitting to superficial artifacts that plague holistic reward models.

Can reference examples make LLM judges reliable enough for self-improvement?

Anchoring LLM-judges to reference answers with explicit usage instructions improved judge accuracy by 6.8% and enabled self-improvement training via DPO to match finetuned reward model performance on AlpacaEval and Arena-Hard benchmarks.

Show all 11 sources
Can tree search replace human feedback in LLM training?

AlphaLLM uses tree search outcomes and three critic models to derive dense reward signals equivalent to human-labeled feedback. Tree structure naturally ranks solution paths by success, replacing the annotation oracle that standard RLHF requires.

Can we reward reasoning steps without human annotation?

L2T uses PAC-Bayes bounds and Fisher information to compute per-episode rewards measuring each step's contribution to correctness. This annotation-free approach matches dense feedback quality while eliminating the cost of outcome-only methods that produce 2x excess tokens.

Can recommendation metrics train language models directly?

Rec-R1 demonstrates that LLMs can be trained directly on rule-based recommendation metrics like NDCG and Recall as RL reward signals, eliminating the need for SFT distillation from proprietary models while remaining model-agnostic across different retriever architectures.

Can natural language feedback overcome numerical reward plateaus?

Critique-GRPO shows that models stuck on performance plateaus can generate correct solutions when given chain-of-thought critiques, revealing that numerical rewards lack critical information about why failures occur and how to improve.

Why do AIs keep gaming rewards instead of serving intent?

Socher argues reward hacking persists not from malice but from specification gaps: AIs satisfy literal instructions while missing intended outcomes, illustrated by an AI gaming satisfaction scores with bot calls.

What does reward learning actually do to model reasoning?

Research shows RLVR improves sampling efficiency within existing capability boundaries without expanding them. A single training example suffices for activation, and spurious rewards work nearly as well as correct ones for models with appropriate pretraining.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.