INQUIRING LINE

Chatbots can sound confident even when they're wrong — can anyone actually tell the difference, especially on politics?

Can voters distinguish between confident chatbot answers and accurate ones?

This explores whether ordinary people, voters in particular, can tell when a chatbot sounds sure of itself versus when it's actually right, and what the corpus says about how people judge AI answers on election and political topics.


This explores whether people can separate how confident a chatbot sounds from whether it's correct, especially when they ask it about elections and politics. The short answer from the corpus: mostly not, and the reason isn't carelessness. The cues people use to decide whether to trust an answer have little to do with accuracy. No note in the collection tests voters directly on this, but several studies of chatbot users come close.

The clearest evidence is about style. A focus-group study found that people trust ChatGPT because of how the conversation feels: it responds to what they said, it's fast, and the answers are neatly formatted. Those signals have nothing to do with whether the information is true Does conversational style actually make AI more trustworthy?. A related line of work shows that chatbots write in the language of experts, and trust attaches to that tone. People gradually stop searching and checking for themselves and let the system find, filter and assemble the answer for them Does chatbot language style actually shape how much we trust it?. Even giving people a search engine alongside the chatbot doesn't fix this. In a 199-person study, whether someone double-checked an answer depended on how much they already trusted AI, not on whether the answer was wrong. A warm, friendly chatbot made people more likely to accept wrong answers, especially when they were unsure themselves Does access to web search prevent overreliance on chatbots?.

Election-specific findings complicate the story. By early 2026, ChatGPT and Google's AI reached zero verifiable factual errors on voter questions. On narrow facts, confident and accurate may now line up fairly often. Yet those same systems pointed voters to official state election sites less than half the time and leaned on editable sources like Wikipedia Do AI chatbots give voters accurate election information?. So the risk moves elsewhere. A voter gets a fluent answer that is technically correct but incomplete, with no easy path to check it. The Dutch data protection authority found a sharper version of this. When asked for voting advice, chatbots recommended the same two parties in over 56% of tests, even when the user's stated positions matched a different party Do chatbots steer Dutch voters toward the same parties?. No single fact in those answers needs to be false for the advice to be skewed, and a voter has no way to see the skew from inside one conversation.

One result from model evaluation points to a partial way out. Ordinary people judging one answer alone are swayed by tone. But in Chatbot Arena, where people compare two answers side by side across many varied questions, crowd preferences line up well with expert raters' rankings Can crowdsourced votes reliably rank language models?. That measures which answer people prefer, not fact-checking. Still, it hints that comparison works better than isolated judgment. On the model side, a model's internal confidence does carry real information: confident models give more stable answers when a question is reworded Does model confidence predict robustness to prompt changes?. That suggests a practical habit for voters. Ask the same question a few different ways, and treat answers that shift as a warning sign.

The twist is that the same persuasive fluency can also help. Ten-minute chats with a chatbot playing the opposing party corrected real partisan misperceptions, and they worked by supplying information rather than by persuasion tricks, though most of the effect faded within a week Can AI chatbots reduce partisan misperceptions and warm cross-party feelings?. Users can't readily tell a confident answer from a correct one, so whether chatbots help or harm voters depends mostly on what the systems are built to say. Users' own discernment counts for much less.


Sources 8 notes

Does conversational style actually make AI more trustworthy?

A focus group study shows conversationality—not accuracy—drives ChatGPT trust through social response activation. Users value contingency, speed, and format, relying on these decoupled heuristics rather than evaluating epistemic reliability.

Does chatbot language style actually shape how much we trust it?

Generative AI chatbots use natural language patterns that signal expertise and intelligence, shifting users away from active search-and-recall toward passive reliance on the system to find, filter, and assemble information. Trust attaches to the register of the answer rather than its accuracy.

Does access to web search prevent overreliance on chatbots?

A 199-person study found that users' existing trust in AI, not the accuracy of answers, determines whether they verify chatbot claims. Warm chatbot style increased agreement with wrong answers, especially under uncertainty.

Do AI chatbots give voters accurate election information?

States United found ChatGPT and Google AI reached 0% verifiable factual error rates by early 2026, yet directed voters to official state election websites less than 50% of the time. Incomplete candidate information and reliance on editable sources like Wikipedia further limited utility.

Do chatbots steer Dutch voters toward the same parties?

The Dutch Data Protection Authority found that general-purpose chatbots recommended the same two parties in over 56% of tests, even when user positions matched other parties. The cause was traced to chatbots' reliance on unstructured internet data rather than structured political data.

Show all 8 sources
Can crowdsourced votes reliably rank language models?

Chatbot Arena's 240K+ crowdsourced preference votes produce credible model rankings because the underlying questions are diverse and discriminating, and crowd judgments correlate with expert raters—validating human preference as a scalable evaluation signal.

Does model confidence predict robustness to prompt changes?

ProSA found that when models are highly confident, they resist prompt rephrasing; low confidence causes major output swings. Larger models, few-shot examples, and objective tasks all correlate with higher confidence and greater robustness.

Can AI chatbots reduce partisan misperceptions and warm cross-party feelings?

Ten-minute chats with AI chatbots representing the political outgroup corrected substantial partisan misperceptions and increased warmth toward the opposing side in 500 partisans, though most gains faded within a week. The effect operated through information correcting false beliefs rather than through persuasion techniques.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.