SYNTHESIS NOTE
Topics›Psychology Chatbots Conversation›this note

How do users mentally model dialogue agent partners?

Exploring what dimensions matter when people form impressions of machine dialogue partners—and whether competence, human-likeness, and flexibility all play equal roles in shaping user expectations and behavior.

Synthesis note · 2026-02-22 · sourced from Psychology Chatbots Conversation

The Partner Modelling Questionnaire (PMQ) validates a three-factor structure for how users perceive machine dialogue partners. The concept originates in psycholinguistics: people form mental representations of their dialogue partner's communicative and social capabilities, and these representations guide what they say, how they say it, and what tasks they entrust to that partner.

Factor 1 — Communicative competence and dependability (49% variance, α=0.88): Strongest items: competent/incompetent, dependable/unreliable, capable/incapable. This is the largest factor — nearly half the variance in how users model a dialogue agent is about whether it can do the job reliably.

Factor 2 — Human-likeness in communication (32% variance, α=0.80): Strongest items: human-like/machine-like, life-like/tool-like, warm/cold. Humans act as the archetype for evaluating communication partners. Even when using machines, people evaluate against a human standard.

Factor 3 — Communicative flexibility (19% variance, α=0.72): Items: flexible/inflexible, interactive/stop-start, interpretive/literal, spontaneous/predetermined. This factor captures whether the agent feels like a living conversation or a scripted interaction.

The definition of partner models — "an interlocutor's cognitive representation of beliefs about their dialogue partner's communicative ability, multidimensional, initially informed by experience and stereotypes, dynamically updated during dialogue" — positions this as the HCI equivalent of theory of mind. Users build, maintain, and update these models continuously.

The practical implication: designing for perceived competence matters most (49% of variance), but human-likeness and flexibility are not negligible. An agent that is reliable but inflexible and machine-like will be perceived very differently from one that is reliable, warm, and spontaneous.

Inquiring lines that read this note 94

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How do agents learn to distinguish valuable feedback from noise? Why do confident AI outputs mislead human trust calibration? Can AI systems participate in genuine communication or only simulate it? How should AI agents balance proactive engagement with conversational respect? How do philosophical assumptions about AI consciousness affect practical harms and design? Do persona-based approaches introduce systematic biases in user simulation? Why do autonomous agents misreport success on failed actions? What enables conversational agents to guide rather than just respond? How do network effects and self-selection distort aggregated rating accuracy? Can confidence signals reliably detect flawed reasoning in language models? Does AI deployment reduce or exacerbate workplace inequality and income instability? What design features sustain romantic bonds with AI companion systems? Do language models reason through disagreement or only accommodate it? How does personalization simultaneously affect user trust and privacy concerns? Why don't better reasoning capabilities improve theory of mind performance? Can AI chatbots provide mental health support without reinforcing harmful beliefs? How should human-AI contributions be measured, disclosed, and verified? Can smaller specialized models match frontier models on key metrics? When do multi-agent systems improve over single frontier models? How can agents discover and adapt to user preferences during conversation? How does AI adoption reshape collaboration patterns in knowledge work? How do users confuse explanation quality with actual system accuracy? Can language models reliably simulate personas and predict behavior? How can AI systems maintain consistent personas across conversations? What distinguishes genuine communicative competence from surface language performance? What determines AI's persuasive power and how can it be detected or mitigated? Does AI assistance help or harm professional skill development? Do single-axis benchmarks accurately measure agent capability for real deployment? How should humans and AI agents share control and decision-making? How susceptible are language models to conversational persuasion and belief change? Why do people trust AI chatbots with sensitive information? Does AI assistance erode cognitive skills while inflating perceived competence? How can emotionally responsive AI maintain reliability and healthy boundaries? How do clinicians calibrate trust in AI medical recommendations? Should GUI agents use structured screen representations instead of end-to-end vision? Are AI-generated articles systematically disadvantaged in search ranking and user engagement?

Related concepts in this collection 2

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
13 direct connections · 115 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

partner models for dialogue agents decompose into three factors — communicative competence human-likeness and communicative flexibility