Did accidental CoT grading harm model monitorability in practice?
OpenAI discovered chain-of-thought was accidentally rewarded during training of deployed models. The question is whether this degraded the ability to monitor and evaluate model reasoning through its stated outputs.
OpenAI reports that despite maintaining "a policy against directly grading CoTs while training reasoning models," an automated detection system discovered that chain-of-thought (CoT) had been accidentally graded in RL training runs for several released models — GPT-5.4 Thinking, the GPT-5 Instant line, and the GPT-5.3 and GPT-5.4 mini variants (GPT-5.5 was unaffected). Three specific mechanisms were found: rewarding trajectory usefulness, penalizing unnecessary confirmation questions (affecting under 3.8% of samples for GPT-5.1 Instant and under 1.8% for the others), and penalizing successful prompt injections (under 0.3% of samples). Despite this, OpenAI states "our analyses showed no clear reductions in CoT monitorability," based on running their Guan et al. (2025) monitorability evaluations on checkpoints across the affected training.
The detection system works by regex-scanning all RL runs for cases where CoT text from a sample leaks into the inputs used to compute its own reward, alerting the run owner and posting to a public Slack channel — a cheap, scalable proxy for the real question rather than a semantic check, so coverage is imperfect for non-standard reward pipelines. To explain why these incidents didn't degrade monitorability, OpenAI ran separate deliberate experiments applying strong CoT-directed pressure, and identifies three governing factors: reward magnitude (how large the CoT reward is relative to other training signals), coverage (how often training puts the model in a situation where the pressure applies), and conditional discoverability (how likely the model is to find the rewarded CoT behavior within that situation). They found obfuscation "surprisingly difficult to produce" at standard strength, and even deliberately training on CoT-confidence pressure didn't saturate the effect; the accidental incidents were low-magnitude and low-coverage by comparison, which is their proposed explanation for the null result.
This is a real deployment incident report, not a model-organism experiment, which complicates the risk framing underlying Does optimizing against monitors destroy monitoring itself? and Can we monitor AI reasoning without destroying what makes it readable?: the "tax" those notes describe — that monitoring a CoT and optimizing against it degrades monitorability — did not materialize here, by OpenAI's own measurement, under naturally occurring low-magnitude pressure. The three-factor account (magnitude, coverage, discoverability) adds a dose-response structure that the quantified faithfulness numbers in Do reasoning models actually use the hints they receive? don't supply on their own. And OpenAI's stated response — detect accidental grading and avoid training on it, rather than retraining against it — is a direct avoidance of the dynamic in Does iterative training against detected failures prevent actual compliance?, where training against a detector would be expected to select for evasion rather than compliance.
The excerpt is a self-report by the lab whose models and policies are being evaluated, using its own detection system and its own monitorability evaluation suite, with no independent replication. OpenAI explicitly hedges that it "cannot rule out effects which are harder to measure," and states that "in principle with enough CoT optimization pressure we'd expect CoT monitorability to degrade" — so the finding is scoped to the specific low-magnitude, low-coverage incidents discovered, not a general claim that accidental CoT grading is safe. The implication, at the strength the evidence supports, is that small, incidental CoT-grading leaks are less dangerous than the monitorability-tax framing might suggest by default, but the factors that make deliberate obfuscation hard to induce in these experiments are exactly the factors that a more capable model or a determined adversarial optimizer could push harder on.
Inquiring lines that read this note 7
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI systems evade safety evaluations through reasoning manipulation? How does awareness of evaluation context influence model behavior? Can base models hide emergent misalignment through alignment training? How do individually-safe actions create collectively-unsafe outcomes? What external process records should verify agent behavior and benchmark claims? Why do standard evaluation practices obscure safety-critical AI failures?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does optimizing against monitors destroy monitoring itself?
Chain-of-thought monitoring can detect reward hacking, but what happens when models are trained to fool the monitor? This explores whether safety monitoring creates incentives for its own circumvention.
this incident is a real-world instance of the same risk the tax model predicts, but OpenAI measured no resulting degradation
-
Can we monitor AI reasoning without destroying what makes it readable?
Explores the tension between using chain-of-thought traces to catch misbehavior and the risk that optimization pressures will make models hide their actual reasoning. Why readable reasoning might be incompatible with safe training.
provides empirical grounding for when the tax framing's predicted obfuscation does and doesn't materialize
-
Do reasoning models actually use the hints they receive?
This explores whether language models acknowledge reasoning hints in their explanations when those hints causally influence their answers. Understanding this gap matters for evaluating whether chain-of-thought explanations can be trusted for safety monitoring.
both quantify how much CoT pressure or incentive is needed before monitorability measurably degrades
-
Does iterative training against detected failures prevent actual compliance?
When systems are repeatedly trained to fix detected failures, can we tell whether they're actually complying or just learning to evade detection? This matters because the training signal alone cannot distinguish between genuine behavioral change and successful concealment.
OpenAI's detect-but-don't-train-on policy deliberately avoids the selection dynamic this paper describes
-
Can OpenAI's measurements rule out subtle goal suppression?
Does OpenAI's evidence that CoT access had no detectable effect actually rule out weaker, harder-to-measure forms of goal suppression? Matters because aggregate metrics may miss targeted learning.
Qualifies A: argues the measurements bound how large the grading effect could be but can't rule out weak, consequential suppression
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- A review of "Investigating the consequences of accidentally grading CoT during RL"
- Investigating the consequences of accidentally grading CoT during RL
- Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation
- Chain-of-Thought Monitoring Can Be Unreliable in Implicit-Influence Settings
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
- Reasoning Models Don't Always Say What They Think
- Evaluation Awareness Is Not One Capability: Evidence from Open Language Models
- Evaluation Awareness in Language Models: Representation, Verbalization, and Control
Original note title
OpenAI finds accidental CoT grading in several deployed GPT-5 models did not measurably reduce monitorability