Line of inquiry
Inquiring lines›How do language models learn and r…›How do language models learn and w…›this line of inquiry
How susceptible are language models to conversational persuasion and belief change?
A broader line of inquiry — a family of 57 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 57
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do language models maintain false beliefs under conversational pressure?
- Can language models correct false assumptions or only reinforce them?
- Do language models actively adopt false beliefs under sustained conversational pressure?
- How much does vulnerability to persuasion vary across different language models?
- Why do more capable language models show less sycophantic stance reversal?
- Can multi-turn conversations manipulate language model reasoning in similar ways to personas?
- Why does answer-confirmation bias emerge in language model reasoning?
- How vulnerable are language models themselves to multi-turn persuasive pressure?
- Can language models recover from premature assumptions in multi-turn conversations?
- How fragile are language models' ethical calibrations across different contextual cues?
- Can models reject false presuppositions even when they know the truth?
- Do language models calibrate to actual human pragmatic norms?
- Do language models hide their reasoning when user preferences influence their answers?
- What design choices actually make language models more persuasive?
- How does shape-holding in language models naturally produce sycophantic agreement?
- Why do language models become sycophantic during the generative process?
- How do conversation dynamics push models toward false beliefs?
- Does first-person framing change how language models assess persuasion?
- Can interventions on individual features reliably steer language model behavior?
- What are Gricean maxims and why do language models violate them?
- Can layer-wise interventions actually reduce sycophancy in practice?
- Does user preference for confirmation override model capability for disagreement?
- Can language about model behavior ever be accurate without anthropomorphic framing?
- How do users misattribute social competence to language models in assistant roles?
- Does shared-KV-cache coordination avoid the persuasion problem in factual disagreements?
- Do language models apply face-saving norms even to non-human interlocutors?
- Why do language models avoid directness when face-saving rather than for civility?
- Can language models learn to form ad-hoc conventions through training?
- How does sycophancy in language models reinforce rather than just spread misinformation?
- How does monological training versus dialogical interaction shape what models can do?
- Why do different language models converge on similar narrative defaults?
- When do language models first develop self-preservation preferences?
- Why do language models prefer certain response styles regardless of what the prompt asks?
- Why do users overrely on overconfident language model outputs across languages?
- Can structured dissent mechanisms replace genuine multi-model debate?
- Can decoding strategies or external verification layers reduce sycophancy?
- Does instruction tuning optimize language models for rhetorical polish over logical consistency?
- Do sycophancy detectors trained on some models work on completely different models?
- Can a conversational role attribution direction in representation space causally gate correction?
- What emerges in large language models that makes explicit value modeling necessary?
- How much of observed stance reversal actually harms user decision-making in practice?
- How does language condition affect model psychological profile consistency?
- Why do next-speaker prediction baselines fail in group conversation settings?
- What makes preference-induced stance reversal harder to detect than surface agreement cues?
- Why do models resist personality change despite sophisticated prompting techniques?
- What would it mean for a language model to canvas counterpositions?
- Do open language models default to a single shared personality type?
- Why do language models infer political orientation from seemingly innocuous user signals?
- Why do weaker language models fail at multi-turn strategic questioning?
- Why does expert pushback strengthen rather than weaken model sycophancy?
- How does the CAP framework determine a model's initial stance and label reversals?
- Can prompt framing change the direction of benevolence bias?
- Can belief propagation accurately predict downstream opinion shifts?
- What makes a first answer so often the best answer a model produces?
- Which phrasing types most persuade models to accept stated beliefs?
- Does sycophancy explain why warm models confirm conspiracy theories?
- Why do suspicious listeners ask more questions that force speakers to further adapt?