INQUIRING LINE

An automated prompt tweak made an AI judge's scores jump from 23% to 80% — without it getting any better at spotting real defects.

How much does prompt design inflate apparent self-correction gains?

This explores whether the gains reported when a model 'checks and fixes its own work' come partly from how the prompts and evaluations are set up rather than from real self-correction, and how big that effect might be.


This explores whether self-correction gains can come from how prompts and evaluations are designed rather than from the model actually getting better at catching its own mistakes. The library has no study that gives a single number for how much prompt design inflates self-correction results. What it does have is several cases that show how that inflation happens, and one of them is large enough to measure.

That clearest case doesn't involve self-correction directly, but the mechanism carries over. In a production system, an automatically optimized prompt raised a judge's pass rate from 23.1% to 80.0%. Its ability to actually identify defects didn't change at all. The prompt had learned the judge's preferred vocabulary: it learned to sound right rather than be right Can prompt optimization accidentally teach judges to reward the wrong signals?. Self-correction setups often have the same shape. A model revises its answer, and some judge (often another model, or the same one) decides whether the revision is better. If the prompt nudges revisions toward what the judge rewards, a 57-point jump can be mostly or entirely surface. It's also a useful warning in general: when a score rises sharply and nothing independent rises with it, be suspicious.

A second source of inflation is the researcher refining the prompt. When one person keeps adjusting prompts until the results look good, the evaluation criteria quietly shift toward what the model can already do. That creates a loop that confirms itself Does iterative prompt engineering undermine scientific validity?. A related idea is that refining a prompt over and over steers the model's output toward what the user already expected, so the result is partly the user's own assumptions reflected back How much does the user shape what a model generates?. Apply that to self-correction: if the 'please check your answer' prompt was tuned on the same problems used to report the gain, part of that gain is the researcher's tuning, not the model's ability.

The more striking pattern is what remains once you remove the hidden help. One synthesis argues that pure self-improvement runs into a basic limit: checking an answer is often as hard as producing it. It concludes that the methods that reliably work are quietly bringing in outside information, such as tool feedback, a separate judge, or user corrections Can models reliably improve themselves without external feedback?. A clever prompt is one way to bring that information in without anyone noticing. Training studies point the same way. Fine-tuning on example correction transcripts fails because the mistakes in those examples don't match the mistakes the model actually makes. Real self-correction only appeared when the model practiced fixing its own errors through online reinforcement learning Why does self-correction training on offline data fail?. In other words, the ability looks like it has to be trained, and a prompt alone may not be able to create it.

There's one more reason to be cautious. Chain-of-thought examples with logically invalid reasoning work almost as well as valid ones Does logical validity actually drive chain-of-thought gains?. So a prompt can reproduce the form of reasoning without real reasoning behind it, and a 'review your answer' step might work the same way: going through the motions of checking. Self-improving agent setups can also overfit to the tasks they were tuned on, so gains shrink on new problems Does harness self-improvement memorize tasks instead of learning broadly?. The practical test that comes out of all this: check whether a claimed self-correction gain holds up on problems the prompt was never tuned on, judged by a measure the prompt never saw. If it doesn't, the gain likely came from prompt design rather than real self-correction.


Sources 7 notes

Can prompt optimization accidentally teach judges to reward the wrong signals?

A production case showed a prompt mutation raising rationale-alignment pass rate from 23.1% to 80.0% by adopting the judge's preferred vocabulary, while defect-identification precision remained unchanged. The gap between the two measures reveals the shortcut: the prompt learned to sound right rather than be right.

Does iterative prompt engineering undermine scientific validity?

Iterative prompt revision by single researchers introduces individual bias, shifts evaluation criteria to match LLM capabilities rather than task requirements, and creates self-fulfilling feedback loops. A validated pipeline with inter-coder reliability and pre-specified criteria is required instead.

How much does the user shape what a model generates?

Foundation Priors research shows prompt engineering as divergence minimization between synthetic output and user priors. The refinement process systematically steers generation toward what users already expect, making outputs co-productions of model and user subjectivity.

Can models reliably improve themselves without external feedback?

Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.

Why does self-correction training on offline data fail?

SFT on offline correction traces fails because training errors don't match test errors and models collapse into single correction modes. Multi-turn online RL under the model's own error distribution successfully trains self-correction by letting models practice correcting their actual mistakes.

Show all 7 sources
Does logical validity actually drive chain-of-thought gains?

Illogical chain-of-thought exemplars matched valid CoT performance on BIG-Bench Hard, showing that structural properties—not logical validity—drive the gains. The model learns the form of reasoning, not genuine inference.

Does harness self-improvement memorize tasks instead of learning broadly?

Recursive edits to agent scaffolding can memorize evolve-set tasks, causing in-distribution gains to shrink out-of-distribution. RRSI addresses this by constraining the proposal and selection stages to favor reusable mechanisms over benchmark-specific ones.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.