SYNTHESIS NOTE
Topics›Personas Personality›this note

Can imaginary listeners reduce dialogue agent contradictions?

Does simulating how an imaginary listener would interpret an utterance help dialogue agents maintain persona consistency without extra training? This explores whether pragmatic self-monitoring at generation time can replace costly supervised approaches.

Synthesis note · 2026-04-18 · sourced from Personas Personality

Persona-based dialogue agents routinely contradict their own stated attributes. Previous solutions either require Natural Language Inference (NLI) labels for training or attach extra trained modules. The "Will I Sound Like Me?" approach (2020) takes a different path: it endows existing agents with public self-consciousness at inference time through an imaginary listener, inspired by social cognition and pragmatics.

The mechanism uses the Rational Speech Acts (RSA) framework. Before generating an utterance, the agent simulates how a listener would interpret it — specifically, whether the listener could distinguish this speaker's persona from a distractor persona based on the utterance. Utterances that would not help identify the speaker (because they are generic or contradictory) are suppressed. The agent learns to ask: "Would I sound like me if I said this?"

The framework extends beyond persona to context consistency in general dialogue. The distractor selection — which alternative persona to contrast against — can be learned rather than manual or random.

This connects to Why does supervised learning fail to enforce persona consistency? as a complementary approach: offline RL punishes contradiction during training while RSA prevents it during inference. The RSA approach requires no additional training data but operates at generation time, adding computational cost per utterance. The trade-off is training-time correction vs. inference-time self-monitoring — both address the same root cause that generative models are never explicitly rewarded for consistency.

The deeper insight is that persona consistency is fundamentally a pragmatic property, not a semantic one. It is about how utterances function in identifying a speaker, not just about logical compatibility of facts.

Inquiring lines that read this note 37

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Do persona-based approaches introduce systematic biases in user simulation? What distinguishes genuine communicative competence from surface language performance? How can AI systems maintain consistent personas across conversations? Can language models reliably simulate personas and predict behavior? Can AI systems participate in genuine communication or only simulate it? Is embodied interaction necessary for language meaning and agency? What structural patterns sustain successful multi-turn dialogue and prevent breakdown? Which reinforcement learning modifications most improve dialogue quality in language models? How should recommendation systems balance individual preference and diversity? Can language models reason beyond surface pattern matching? What prediction granularity best trains models to generate reliable reasoning? How susceptible are language models to conversational persuasion and belief change? Can readers reliably distinguish AI-written text from human writing?

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

pragmatic self-consciousness through an imaginary listener reduces persona contradiction without additional training — Rational Speech Acts framework enforces consistency by simulating how utterances would be interpreted