SYNTHESIS NOTE
Topics›Conversation Architecture Structure›this note

Can models learn to abstain when uncertain about predictions?

Explores whether language models can be trained to recognize when they lack sufficient information to forecast conversation outcomes, rather than forcing uncertain predictions into confident-sounding responses.

Synthesis note · 2026-02-22 · sourced from Conversation Architecture Structure

Generating a single plausible next-utterance is not the same as modeling the uncertainty about ALL possible next-utterances in a calibrated way. In negotiations, "Sounds good!" and "No thanks" may be equally fluent/topical/informative responses, but one may be more likely given the goals, beliefs, and emotions of the interlocutors.

FortUne Dial formalizes this as conversation uncertainty modeling, shifting evaluation from pure accuracy to uncertainty-aware metrics that enable abstention on individual instances. When the model estimates high uncertainty about an outcome, it should say "I don't know" rather than forcing a prediction.

Two representations of uncertainty:

Two fine-tuning strategies improve calibration:

The practical result: smaller open-source models, once calibrated, can compete with pre-trained models 10x their size on uncertainty-aware forecasting. This suggests that calibration ability is undertrained in standard LLMs — the capability exists but the training signal is absent.

Applications include: studying effects of strategy and social structure in negotiations, intervening to improve human and machine conversations, and assessing trust/heterogeneity in data sources via entropy metrics.

Real-world deployment evidence from CRAFT: When the CRAFT conversational forecasting model was deployed as a prototype moderation tool for Wikipedia editors, moderator feedback revealed critical design dimensions. Score change (trajectory) was more actionable than absolute score — moderators preferred seeing whether a conversation was trending toward derailment rather than a static risk number. Crucially, moderator confidence in predicting derailment varied dramatically: four of nine participants believed they could forecast in any Wikipedia context, four others only in very specific contexts with low confidence, and one only for personally-known participants on familiar topics. This variance means forecasting tools must accommodate heterogeneous human expertise rather than assuming uniform detection ability. A further missing dimension: conversation age. Moderators reported that inactive conversations (>2-3 days since last comment) are unlikely to revive, much less turn uncivil — but the prototype did not surface this temporal signal. The scale problem is stark: even topic-engaged moderators cannot proactively monitor all at-risk conversations, forcing them to rely on random discovery strategies.

Since Does reasoning fine-tuning make models worse at declining to answer?, calibrated uncertainty and appropriate abstention are capabilities that current training actively degrades. Since Does training objective determine which direction models fail at abstention?, the direction of calibration failure depends on the training regime — a forecasting system built on reasoning-trained models would over-predict, while one built on safety-trained models would refuse to predict. Conversation forecasting requires the opposite of both failure modes: models that know what they don't know about where a conversation is heading.

Additional empirical domain — Instagram hostility forecasting: A separate forecasting study on Instagram demonstrates that hostile comments can be predicted from early conversational signals: AUC 0.82 for predicting hostility presence 10+ hours in the future, and AUC 0.91 for predicting whether a post will receive more than 10 hostile comments vs. only one. Predictive features include the post author's history of receiving hostile comments, user-directed profanity, number of distinct participants, and hostility trends in the conversation so far. This complements the CRAFT deployment evidence above — different platform, similar principle: early conversational dynamics carry forecastable signal about future trajectory.

Inquiring lines that read this note 116

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

Which reinforcement learning modifications most improve dialogue quality in language models? What enables conversational agents to guide rather than just respond? What structural patterns sustain successful multi-turn dialogue and prevent breakdown? Why do training associations persist despite contradictory contextual information? What prediction granularity best trains models to generate reliable reasoning? How do users confuse explanation quality with actual system accuracy? Can confidence signals reliably detect flawed reasoning in language models? Can persona profiles improve LLM prediction accuracy and consistency? Should models ask for clarification when facing ambiguous or under-specified information? Why does self-revision amplify confidence in wrong model answers? How susceptible are language models to conversational persuasion and belief change? Why do people trust AI chatbots with sensitive information? What distinguishes genuine communicative competence from surface language performance? What structural biases does transformer attention architecture inherently introduce? Can models develop genuine introspective capability, or only mimic it? How does scaling reasoning capabilities affect models' appropriate abstention behavior? What limits language model accuracy in evaluating ideas? How do reward signal properties affect model reasoning and safety? How should AI agents balance proactive engagement with conversational respect? Can AI chatbots provide mental health support without reinforcing harmful beliefs? Can reasoning models use reflection to correct their initial outputs? Does scaling reasoning capability create fundamental tradeoffs in control and reliability? How can agents discover and adapt to user preferences during conversation? Can artificial systems establish authority in domains requiring expert judgment? Does preference optimization undermine conversational grounding in language models? How do training data quality and composition affect downstream model performance? How does fine-tuning trade off accuracy against reasoning quality? What explains the gap between benchmark scores and true reasoning capability? How can humans maintain effective oversight as AI systems scale? Why do language models struggle to implement user intent accurately from prompts? Can models strategically underperform during evaluation to hide capabilities? How do clinicians calibrate trust in AI medical recommendations? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? How do agents learn to distinguish valuable feedback from noise? How reliably can language models perform causal versus temporal reasoning? What design features sustain romantic bonds with AI companion systems?

Related concepts in this collection 8

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
21 direct connections · 233 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

conversation forecasting under uncertainty requires calibrated probability estimates — calibrated models should abstain on uncertain predictions rather than forcing outputs