Does setting temperature to zero actually make LLM outputs reliable?
Explores whether deterministic LLM settings that produce consistent outputs also guarantee reliable judgments, and how to measure true reliability beyond surface consistency.
"Can You Trust LLM Judgments?" (2024) introduces a rigorous framework for evaluating LLM-as-a-Judge reliability using McDonald's omega, revealing that the common practice of using fixed seeds and deterministic settings provides false confidence.
The core argument: even with deterministic settings, a single LLM output is one sample from the model's probability distribution. Setting temperature to zero and fixing the seed produces "fixed randomness" — the same output every time, but that output may still be a misleading draw from the distribution. Consistent replication does not guarantee reliability. A perfectly calibrated LLM that says it's 90% confident should be correct 9 out of 10 times — but even a perfectly calibrated LLM can be unreliable if its distribution has high variance.
The framework: prompt the judgment LLM 100 times, varying only the replication while holding all other factors constant. Apply McDonald's omega to assess internal consistency across these replications. This reveals whether the model's judgments are stable properties of the input or artifacts of the sampling process.
The distinction between reliability, confidence, and calibration is critical:
- Calibration: alignment between stated confidence and actual correctness
- Confidence: the model's self-assessed certainty
- Reliability: consistency of judgments across multiple draws
These three are intertwined but distinct. A model can be well-calibrated (confident when right) but unreliable (different answers on different draws). A model can be reliable (always gives the same answer) but poorly calibrated (that consistent answer is wrong).
This connects to Does model confidence predict robustness to prompt changes? — ProSA measures sensitivity to prompt variation, while this measures sensitivity to sampling variation. Both reveal that single evaluations are insufficient. The practical implication: any LLM-as-a-Judge deployment that relies on single-shot evaluation with deterministic settings is providing the illusion of precision without evidence of reliability.
Inquiring lines that read this note 191
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Why do language models hallucinate and how can we prevent it?- What makes LLM outputs fabrication rather than hallucination or confabulation?
- What should we call errors in LLM outputs when hallucination does not apply?
- How much does ROUGE metric choice inflate hallucination detection claims?
- Does inevitable LLM hallucination make detection metric validity critical?
- Can LLMs evaluate their own observations without external feedback?
- Can systems lacking inner states express genuine truthfulness claims?
- When does provable stability in latent dynamics fail to preserve fidelity?
- Why do one-shot transparency studies miss the temporal reversal entirely?
- Can a correct outcome hide a fundamentally unsound decision-making process?
- How can a single instrument measure errors across multiple system layers?
- Why does aggregate accuracy fail as a metric for rare harmful cases?
- How does unidimensionality in assessments affect measurement validity?
- Why do standard accuracy metrics ignore set-level consumption constraints?
- What makes the 45 percent accuracy saturation threshold universal?
- Why does sophisticated measurement not validate the underlying scientific inference?
- Can similar outputs from different systems prove they work the same way?
- How should benchmarks balance verifiability against outcome resolution?
- How can hidden test partitions detect constant predictions that generalize?
- How would longitudinal measurement reveal sustained effects of friction?
- Why do cumulative-best reporting inflate progress compared to actual validation?
- How does the absence of failure rate information affect generalizability claims?
- What distinguishes minimal-pair asymmetry from standard accuracy evaluation?
- What makes the Brier score mathematically better than log-likelihood here?
- Can utility control modify LLM values more effectively than output filtering?
- Can scaling predictions become reliable if improvements are continuous not sudden?
- How do surface statistical regularities enable correct outputs while degrading robustness?
- What makes output convergence across models inevitable given input-side homogenization?
- What consumption data would validate the limited-consumption model in production systems?
- Why do rare cases in medicine and science require models that preserve tail distributions?
- Can deterministic computation actually create new information in data?
- Can population-level distributions shift usefully even when individual prediction fails?
- Can imperfect uncertainty estimates still beat uniform oversight strategies?
- Can experimental outcomes be reliably distilled into reusable insights?
- Can aggregate survey realism coexist with unreliable fine-grained effects?
- Can defenders tighten the total-variation bound in practice with measured benign activation rates?
- What makes a bounded observer's ability to extract information different from apparent randomness?
- What would it mean to assign explicit trust weights to synthetic data?
- What role does real-time accuracy feedback play in reducing user overreliance?
- Can LLM judges reliably estimate when they lack sufficient persona information?
- What does the 20-questions test reveal about LLM character consistency?
- Can LLMs recover true joint distributions from marginal census data?
- Can we detect superposition in LLM personality traits and stated preferences?
- How can we validate LLM-based drift measurements against human judgment?
- Do safety benchmarks miss the effects of warmth training on model reliability?
- Can safety benchmarks detect reliability degradation from warmth training?
- What calibration corrections can reduce LLM judge bias in automated evaluation pipelines?
- What does McDonald's omega reveal about LLM judgment consistency?
- How do calibration and reliability differ in LLM judge evaluations?
- What other evaluation biases exist in LLM judge systems?
- Can crowdsourced voting and automated panels both credibly evaluate LLM outputs?
- Where should measurement systems sit to avoid recording bias?
- How can deterministic checks make wrong judge decisions survivable?
- What makes a judge's calibration at decision boundaries harder to improve?
- How sensitive are LLM bias measurements to analysis choices?
- Why do humans and zero-shot LLM judges perform worse than trained detectors?
- Why do LLM-judged tournaments fail to estimate candidate value or uncertainty reliably?
- Why do users systematically overrely on confident LLM outputs across languages?
- Can fact-checking systems use LLMs reliably if models abandon correct positions under pressure?
- Why do true and false LLM outputs use the same mechanism?
- Why do different LLMs converge on nearly identical outputs?
- How do users mistake synthetic LLM outputs for empirical observations?
- What structural features force users to evaluate the epistemic status of outputs?
- When does the correlation between consistency and correctness break down?
- Why does accumulated portfolio output not match accumulated worker capability?
- What makes a deployment paradigm credible for maintaining scientific integrity?
- What makes inter-coder reliability testing essential for prompt validation?
- How can we verify outputs from systems that generate without grounding?
- Which use cases can tolerate unverified LLM outputs without external verification?
- Can lightweight verification methods help experts trust LLM outputs?
- How does step-level confidence filtering compare to global confidence averaging?
- How do we assign confidence and polarity scores to belief edges?
- Does layer-wise prediction stabilization provide a stronger trace quality signal than confidence alone?
- Why does model confidence correlate with robustness to prompt variations?
- Can an LLM be well calibrated but still unreliable on single evaluations?
- How reliable is the top-2 confidence gap as a stopping signal across tasks?
- Can uncertainty estimates based on model self-assessment reliably signal errors?
- Why do improvements in accuracy come at the cost of calibration?
- Does model confidence actually correlate with robustness against prompt variations?
- Why do models confabulate inconsistently across different samples?
- Can semantic entropy improve model calibration without external ground truth?
- Can proper scoring rules restore model calibration without sacrificing accuracy?
- What makes mathematically confident but incorrect answers resemble valid solution shapes?
- How does confidence in LLM outputs override users' ability to check accuracy?
- How can distillation preserve uncertainty expression instead of optimizing it away?
- How do local soundness signals work across different problem domains?
- What makes uncertainty calibration harder than expanding knowledge?
- Can log-probability confidence be separated from decision-aligned signals?
- How do confident system outputs weaken user skepticism about their reliability?
- Why do fluent predictions fail to capture reliable internal models?
- What makes confident hallucination a distinct problem from poor calibration?
- Why is faithful calibration considered fundamentally metacognitive?
- Can retrieval density alone correct overreliance on confident but wrong outputs?
- How many samples are needed to distinguish systematic error from genuine uncertainty?
- Does treating model disagreement as belief rather than noise change how we audit outputs?
- How should product specifications measure alignment without naming the dimension?
- What happens when alignment targets measure only the preferred dimension of entangled properties?
- Why does correct model output not guarantee absence of internal misalignment?
- Can distributional views explain when an LLM appears to change its mind?
- Why does regenerating LLM responses produce different but equally valid answers?
- Can LLMs express uncertainty in ways that preserve epistemic honesty?
- Can evaluation criteria be reliably encoded in labeled data without ground truth standards?
- How should process quality and verification cost factor into evaluation judgment?
- Does majority voting reliably signal correctness without risking reward hacking?
- Why does low temperature sampling extract consensus from diverse training data?
- What signals detect when consensus training is silently degrading performance?
- Can log-likelihood loss combined with binary rewards achieve calibration?
- How does self-consistency compare to confidence as a proxy reward signal?
- How does 93% reward reliability compare to other RL noise sources?
- How does activation consistency training differ from output-level consistency?
- Why does test accuracy improve after training accuracy reaches 100 percent?
- Can model updates be designed to prevent simultaneous acceptance of conflicting facts?
- Why does analytical depth demand trigger fabrication over transparent uncertainty?
- What property must remain constant to individuate an LLM across infrastructure changes?
- Does exposure to more domain-specific examples reduce LLM overconfidence?
- What makes some model capabilities reliable while others remain brittle?
- Can users experience the LLM Fallacy even when AI outputs are completely accurate?
- What happens when we treat LLM outputs as sampled rather than stored?
- Can we systematically enumerate LLM failure modes from first principles?
- Do longer prediction horizons systematically degrade LLM forecasting accuracy?
- Why do LLM outputs need verification even when they look polished?
- How do training data cutoffs produce false claims that stay consistent?
- Why do models fail under distribution shift if accuracy metrics stay high?
- What baseline evidence distinguishes amplification from unchanged failure rates?
- How should monitoring intensity change based on task criticality?
- Can per-decision human review ever maintain capacity against volume and fatigue?
- What distinguishes actual social disagreement from distributional uncertainty in LLM outputs?
- What structural coherence exists in LLM preference systems and value hierarchies?
- Can measuring semantic entropy help us detect unreliable generations?
- Can contamination-free evaluation distinguish between memorization and genuine prediction ability?
- How much does omniscient evaluation overstate real-world simulation fidelity?
- Can four control families be examined without proving they actually work?
- What distinguishes authentic consistency from Hawthorne-effect confounds in benchmarks?
- Why is the Judging preference constant while other traits vary slightly?
- Do high-disagreement items signal contested values or measurement noise?
- What consistency tests could distinguish constructed from genuine preferences?
- Why does preference measurement validity matter more than aggregation methods?
- Why does preference measurement validity matter before any aggregation?
- Does measured opinion stability reflect genuine preference or elicitation artifact?
- How does Goodhart's Law apply when safety measures become optimization targets?
- What makes uniform bounds the right choice for safety boundaries?
- Why does self-consistency fail as a proxy reward for correctness?
- Can models detect statistical properties of their own generation in real time?
- What makes self-consistency a sufficient training target for the judge role?
- How does distilling only inconsistent rollouts compare to distilling all generations?
- Can self-ratings of output quality predict forecast performance?
- Does higher temperature sampling help bootstrapping loops generate more training data?
- What specific bookkeeping tasks can environments maintain more reliably than policies?
- Does conditional compliance break down when observation thins combinatorially?
- Can per-user adapters remain consistent without drifting or leaking?
- How much do different LLMs independently converge on similar outputs?
- What trade-offs emerge between training objectives and model reliability?
- Why is measurement during training critical before deployment?
- Can trustworthy scoring prevent persistent iteration from compounding errors?
- Can a metric that rewards central tendency hide degenerate predictor failures?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does model confidence predict robustness to prompt changes?
Explores whether a model's certainty about its answer determines how much it resists prompt rephrasing and semantic variation. This matters because it could explain why some tasks are harder to evaluate reliably.
prompt sensitivity and sampling sensitivity are complementary reliability concerns
-
Can LLM judges be fooled by fake credentials and formatting?
Explores whether language models evaluating text fall for authority signals and visual presentation unrelated to actual content quality, and whether these weaknesses can be exploited without deep model knowledge.
judge unreliability compounds with exploitable biases
-
Why do preference models favor surface features over substance?
Preference models show systematic bias toward length, structure, jargon, sycophancy, and vagueness—features humans actively dislike. Understanding this 40% divergence reveals whether it stems from training data artifacts or architectural constraints.
calibration failure at the preference model level adds to the reliability problem
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
- Can Large Reasoning Models Self-Train?
- LLM-as-a-Judge Is Not an Oracle: Why Self-Improving Agents Need Deterministic Guardrails
- AbstentionBench: Reasoning LLMs Fail on Unanswerable Questions
- Using Large Language Models to Create AI Personas for Replication and Prediction of Media Effects: An Empirical Test of 133 Published Experimental Research Findings
- Can Machines Think Like Humans? A Behavioral Evaluation of LLM-Agents in Dictator Games
- DecepChain: Inducing Deceptive Reasoning in Large Language Models
- When Hindsight is Not 20/20: Testing Limits on Reflective Thinking in Large Language Models
Original note title
deterministic LLM settings create fixed randomness not reliability — a single output remains one draw from the model's probability distribution