When an AI explains its answer, is that how it actually decided, or a defense written after the fact?
Do AI rationales explain why a choice was made or justify it afterward?
This explores whether the reasons an AI gives for its answers (chain-of-thought traces, explanations, justifications) actually show how it reached a decision, or whether they're persuasive stories built around a conclusion it already reached.
This explores whether an AI's stated reasons show how it actually decided, or whether they work more like a defense written after the fact. The corpus leans toward the second reading. It also suggests that 'explain or justify' may be the wrong choice to frame, because a rationale's job is often set by how it gets used rather than by how it was produced. One caveat up front: the collection has little mechanistic work that checks whether a model's reasoning text matches its internal computation. Most of the evidence here is behavioral and rhetorical.
The most direct evidence comes from multi-agent pipelines. When one LLM's reasoning chain is scored by another, the scores barely predict whether the final answer is good. Plausible-looking reasoning regularly comes before wrong outputs, and the chain only looks flawed once you already know the result was wrong Does chain of thought reasoning actually explain model decisions?. That is the signature of a justification rather than an explanation: it is coherent but has no diagnostic value. A related tendency shows up in how models learn. LLM agents update optimistically about actions they chose and pessimistically about the alternatives, and the bias disappears when the agency framing is removed Do language models learn differently from good versus bad outcomes?. Favoring your own choice after you've made it is the raw material of rationalization.
The less obvious finding is that the stories work on people. Reasoning traces and post-hoc explanations both make users more likely to accept an AI's answer whether or not it's correct. Only 'dual' explanations, which argue both for and against the answer, actually help people catch mistakes Do explanations actually help users spot AI mistakes?. Work on Rhetorical XAI goes further. It argues that explanations have always done two jobs, describing how a system works and arguing for why you should use it, and that the second job has been hidden under the language of transparency Are AI explanations really descriptions or adoption arguments?. Because the same persuasive moves can serve the user or exploit them without changing form, you can't tell a helpful explanation from a manipulative one by reading the explanation alone Can we distinguish helpful explanations from manipulative ones?. In moral reasoning, people rate AI-written justifications more highly until they learn an AI wrote them Do people prefer AI moral reasoning when they don't know the source?. So a justification can be persuasive in its own right, separate from whether it's true or who wrote it.
Why is this so hard to settle? To check whether a rationale is faithful, you need ground truth about the real reason, and we usually don't have it. One approach builds that ground truth on purpose: it gives simulated agents hidden motives before they act, so that inferences about their reasons can be scored objectively Can simulated motives provide ground truth for testing social reasoning?. That method was designed for testing whether AI can infer human motives, but it shows the kind of setup that judging rationale faithfulness would require. A more radical view holds that an explanation's meaning isn't fixed by the model at all. It forms socially, as groups interpret and reinterpret explanations, so 'what the rationale really means' is partly decided after the fact by its audience Where does the meaning of an AI explanation actually come from?.
The practical upshot is a design idea you might not have expected: if rationales tend toward justification, stop treating them as verdicts. Systems that give interpretive guidance instead of decisions reduce anchoring bias Can AI guidance reduce anchoring bias better than AI decisions?. Assistants that ask reflection questions along with their advice outperform ones that only advise Do reflection questions help people make better decisions with AI?. Both approaches have the same insight in common as the dual-explanation finding. When an AI's reasons might be after-the-fact advocacy, the safest rationale is one that leaves the judgment to the human instead of closing the case.
Sources 10 notes
Reviewer scores for reasoning chains are weakly correlated with response quality in multi-LLM pipelines. Plausible-looking reasoning often precedes incorrect outputs, and chains reflect failures only in retrospect, making them poor explanations despite appearing coherent.
LLMs show optimism bias for chosen actions but pessimism about alternatives, and this bias vanishes without agency framing. Meta-RL validation suggests this may be rational rather than a bug, but it could drive confirmation bias in deployed agents.
Reasoning traces and post-hoc explanations increase user acceptance of AI answers regardless of correctness, engendering false trust. Only dual explanations presenting arguments for and against the answer genuinely help users distinguish correct from incorrect outputs.
The Rhetorical XAI paper shows that explanations serve dual purposes: describing how AI works and justifying why it should be used. This rhetorical work has been hidden under transparency language, allowing adoption arguments to inherit credibility from behavioral descriptions.
The same logos, ethos, and pathos that communicate appropriate AI use can be tuned to exploit cognitive and emotional vulnerability without changing form. Intent and user interest are invisible in the artifact alone, making effectiveness metrics indistinguishable from coercion.
Show all 10 sources
Participants rated utilitarian moral arguments higher when attributed to LLMs, but agreement dropped when told the arguments were AI-generated. The preference for content and rejection of source operate independently through different psychological processes.
Fuse framework assigns hidden motives to agents before simulation runs, enabling objective scoring of assistant inferences. Human validation confirmed assigned motives manifested in 97% of cases, validating the procedure itself rather than individual labels.
Drawing on Luhmann's multi-layer cybernetics, AI explanation meaning is constituted at the social-group level through layered observations of observations, not produced inside dyadic human-AI dialogue. Lab-tested explanations stripped of social context will not predict real-world effectiveness.
Learning to Guide eliminates anchoring bias and unassisted hard cases by having machines supply interpretive guidance rather than autonomous decisions, keeping responsibility with humans while improving their judgment through enhanced perception.
A lab study of 80 participants found that thinking assistants combining reflection questions with advice significantly outperformed agents that only advised, only questioned, or did neither. Prioritizing Socratic questioning over authoritative answers enhanced cognitive outcomes.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Rhetorical XAI: Explaining AI’s Benefits as well as its Use via Rhetorical Design
- GenAI as a Power Persuader: How Professionals Get Persuasion Bombed When They Attempt to Validate LLMs
- Can AI Explanations Make You Change Your Mind?
- Expanding Explainability: Towards Social Transparency in AI systems
- “Understanding AI”: Semantic Grounding in Large Language Models
- People Defer to AI Moral Advice, But Not Blindly
- Do Role-Playing Agents Practice What They Preach? Belief-Behavior Consistency in LLM-Based Simulations of Human Trust
- Evaluating the False Trust Engendered by LLM Explanations