INQUIRING LINE

Asking an AI the same question many times shows when it's unsure, but it can't reveal when it's consistently wrong.

How many samples are needed to distinguish systematic error from genuine uncertainty?

This explores whether sampling a model's answer many times can tell you when it is confidently wrong (a repeated, systematic error) versus genuinely unsure, and how many samples that would take.


This explores whether running the same question through a model many times can separate two kinds of failure: the model being unsure (answers scatter) and the model being consistently wrong (it repeats the same mistake). The corpus gives no magic number. What it suggests instead is more useful: adding samples helps with only one of these two failures. Repeated sampling exposes uncertainty well. It cannot expose a systematic error, however many draws you take.

The clearest statement of this is Can agreement across samples reveal when models are wrong?. Checking whether samples agree catches answers that the model makes up differently each time. But when the model repeats the same falsehood every time, agreement looks exactly like confidence. A thousand identical wrong answers look the same as ten. Does setting temperature to zero actually make LLM outputs reliable? makes a related point from the other direction. Setting temperature to zero gives you the same output every time, but that output is still just one draw from the model's distribution. Their reliability testing used 100 repetitions, which is the only concrete sample count in this set, and it was used to measure variability, not to find hidden errors. Consistency and correctness are separate properties.

So the useful question changes from "how many samples?" to "what other signal can break the tie?" The corpus offers several, and they come from different directions. One looks at the data: Can pretraining data statistics detect hallucinations better than model confidence? flags risk by checking whether the entities in a question rarely appeared together in training data. That catches confident hallucinations that sampling would miss. Should RAG systems use model confidence or data rarity to trigger retrieval? shows that the two signals cover different failures. Model uncertainty catches shaky reasoning about familiar topics. Data rarity catches confident errors about unfamiliar ones. Combining them beats using either alone. A second approach tests whether the answer holds up when the question is reworded instead of simply asking it again. Does model confidence predict robustness to prompt changes? finds that answers given with low confidence swing when the prompt is rephrased, which makes rephrasing a cheap probe.

A third approach looks inside each sample instead of counting samples. Does step-level confidence outperform global averaging for trace filtering? shows that tracking confidence at each step of a reasoning trace gets roughly the gains of majority voting with far fewer traces, because it catches the exact point where reasoning breaks down. Can confidence trajectories reveal when reasoning goes wrong? finds a specific pattern behind systematic errors: the model settles on an answer early and then rationalizes it. Confidence that arrives too fast is a warning sign that a simple agreement count would never show. Can we measure how deeply a model actually reasons? goes a level deeper. It measures how much the model's predictions change across its internal layers, and that matches self-consistency at a lower cost.

The point you may not have expected: past a certain point, more samples just confirm the model's blind spots. Telling systematic error apart from uncertainty takes a signal from outside the model's own distribution, such as training-data statistics, reworded prompts, or deterministic checks. Can separating judgment from verification improve research paper reliability? builds a whole system on that idea. If you want a principled way to pick which probes to run rather than how many, Can optimal experimental design improve few-shot example selection? treats example selection as a budget problem: choose the queries that reduce uncertainty the most. That framing could carry over here, but the corpus does not test it for error detection.


Sources 10 notes

Can agreement across samples reveal when models are wrong?

The Consistency Veto suppresses answers that vary across samples but cannot detect systematic errors the model repeats identically. It carries strong signal on some queries but inherits a fundamental blind spot: agreement looks like confidence even when both are wrong.

Does setting temperature to zero actually make LLM outputs reliable?

Fixed seeds and zero temperature replicate the same output repeatedly, but that output remains one draw from the model's probability distribution. McDonald's omega testing across 100 repetitions reveals that consistency does not equal reliability.

Can pretraining data statistics detect hallucinations better than model confidence?

QuCo-RAG uses entity co-occurrence patterns from training data to trigger retrieval, successfully flagging hallucination risk even when models are highly confident. This data-side approach catches the root cause (unseen combinations) rather than the symptom (low confidence).

Should RAG systems use model confidence or data rarity to trigger retrieval?

Model confidence and data-rarity signals catch orthogonal failure modes: confidence misses hallucinations about rare entities, while rarity misses uncertain reasoning about common knowledge. Hybrid triggers substantially outperform either signal alone.

Does model confidence predict robustness to prompt changes?

ProSA found that when models are highly confident, they resist prompt rephrasing; low confidence causes major output swings. Larger models, few-shot examples, and objective tasks all correlate with higher confidence and greater robustness.

Show all 10 sources
Does step-level confidence outperform global averaging for trace filtering?

Local step-level confidence catches reasoning breakdowns that global averaging masks and enables early stopping before traces complete. This approach achieves comparable accuracy gains to naive majority voting with far fewer generated traces, proving trace quality matters more than quantity.

Can confidence trajectories reveal when reasoning goes wrong?

Models that commit to answers early then rationalize show measurable flawed reasoning. Rewarding gradual confidence growth via RL improves accuracy significantly—on Countdown by 42 percentage points—without needing process labels or external reward models.

Can we measure how deeply a model actually reasons?

Deep-thinking ratio (DTR) measures the proportion of tokens whose predictions undergo significant revision across model layers, correlating robustly with accuracy across AIME, HMMT, and GPQA benchmarks. Think@n, a test-time strategy using DTR, matches self-consistency performance while reducing inference costs.

Can separating judgment from verification improve research paper reliability?

Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.

Can optimal experimental design improve few-shot example selection?

AIPD frames demonstration selection as budgeted active learning, choosing examples that maximally reduce test-set uncertainty. Two algorithms (GO and SAL) outperformed similarity-based methods across small, medium, and large language models.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.