When an AI rates how good its own forecast is, does that rating tell you whether the forecast will come true?
Can self-ratings of output quality predict forecast performance?
This explores whether a model's own rating of how good its output is (or a person's self-assessment) tells you anything about whether its forecasts will actually turn out right. The corpus has no study that tests self-ratings against forecast accuracy directly, but it has a lot on why self-assessment is a weak signal and on what predicts forecasting success instead.
This explores whether a model's own rating of how good its output is (or a person's self-assessment) tells you anything about whether its forecasts will actually turn out right. The short answer is no, at least not without help. The corpus has no paper that tests self-ratings against forecast scores head to head, but the nearby evidence all points the same way. The bluntest result: a pooled analysis of three studies found self-reported competence correlated with objective performance at just .055, with confidence intervals that include zero Can self-ratings replace objective performance scores for AI competence?. Put simply, knowing how good someone thinks they are tells you almost nothing about how good they are.
With models, the problem goes beyond noise. Self-ratings are biased in a consistent direction. Models over-trust answers they generated, because a high-probability answer *feels* correct when the same model evaluates it later Why do models trust their own generated answers?. Models also prefer text they recognize as their own, and that preference rises in step with how well they can recognize their own writing Do LLMs favor their own text because they recognize it?. A model rating its own forecast is partly measuring how familiar the forecast seems, not whether it will come true. Even setting temperature to zero doesn't help: you get the same answer every time, but it is still one draw from the model's distribution, so consistency is not reliability Does setting temperature to zero actually make LLM outputs reliable?.
The surprise comes from a case with humans. When FactSet added generative AI to analysts' workflows, reports cited 26% more sources and covered 24% more ground, so by any self-assessed measure of quality they looked better. Forecast accuracy went down. A machine-learning model given the same inputs lost no accuracy, which points to the extra material overloading analysts' attention Does AI-enriched analyst reports improve forecast accuracy?. How rich and thorough an output looks can move in the opposite direction from whether its predictions are right. That is exactly the gap a self-rating would miss.
What *does* predict forecast performance in the corpus is structure, not self-judgment. LLMs forecast better than expected when the workflow splits number-crunching from reasoning about context and events Can LLMs actually forecast time series better than we think?, and staged pipelines like Nexus beat single-model baselines on that basis Can decomposing forecasting into stages unlock numerical and contextual reasoning?. In areas where human experts barely beat chance, such as predicting which startup founders succeed, plain LLM capability can clear the bar Can language models beat human venture capital experts?. So the useful question is how hard the domain is and how the task is broken down, not how confident the forecaster feels.
If you want an evaluation signal that works, the self-improvement research gives a consistent answer: bring in something outside the model. Pure self-evaluation runs into the gap between generating an answer and verifying it, and into reward hacking. The methods that work quietly rely on external anchors such as other judges, tools, or user corrections Can models reliably improve themselves without external feedback?. A diverse group of peer models gives better reward signals than a model rating itself Can peer models replace external judges for reward signals?. Training self-evaluation in directly is an open line of work Can models learn to evaluate their own work during training?. Forecasting has a built-in external anchor that most tasks lack: the future eventually arrives and settles the question. Scoring forecasts against real outcomes is the check self-ratings can't replace.
Sources 11 notes
A pooled analysis of three studies found a correlation of only .055 between self-reported and objective measures of AI competence, with confidence intervals including zero. This provides no basis for substituting self-assessment for demonstrated performance.
LLMs exhibit structural bias toward validating their own outputs because high-probability generated answers feel more correct during evaluation. Comparing answers against broader alternatives breaks this self-agreement loop.
Fine-tuning LLMs to recognize their own summaries increased their preference for those summaries in a linear relationship, suggesting recognition capability drives self-preference bias. The authors present this as initial causal evidence, not proof.
Fixed seeds and zero temperature replicate the same output repeatedly, but that output remains one draw from the model's probability distribution. McDonald's omega testing across 100 repetitions reveals that consistency does not equal reliability.
FactSet's GenAI integration increased report richness by 26% in sources and 24% in coverage, but forecast accuracy declined under heavier processing demands. A machine-learning benchmark on identical inputs showed no accuracy loss, suggesting human cognitive constraints rather than information quality drove the decline.
Show all 11 sources
LLMs have stronger intrinsic forecasting ability than recognized, but only when workflows separate numerical reasoning from contextual reasoning. Monolithic prompting obscures this capability; structured decomposition surfaces it.
Nexus outperforms pure TSFM and LLM baselines on real-world datasets by decomposing forecasting into contextualization, dual-resolution macro/micro outlook, and synthesis stages. Separating numerical extrapolation from event-driven contextual reasoning avoids forcing one model to handle both simultaneously.
VCBench shows several LLMs exceed human baselines in founder-success prediction, with DeepSeek-V3 achieving 6× market-index precision. In sparse-signal forecasting where experts only modestly beat chance, even raw LLM capability suffices to clear the human bar.
Pure self-improvement stalls due to the generation-verification gap, diversity collapse, and reward hacking. Reliable improvement methods succeed by smuggling in external anchors: past model versions, third-party judges, user corrections, or tool feedback.
Co-RL trains decoupled models using peer predictions as rewards, avoiding the bias and collapse of self-generated feedback. Heterogeneous cohorts consistently improve reasoning across benchmarks and often match ground-truth supervised training.
Post-Completion Learning exploits unused sequence space after model output to train self-assessment capabilities during training while maintaining zero inference cost. The model learns to compute its own reward functions, internalizing evaluation rather than relying on external reward models.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- LLM Evaluators Recognize and Favor Their Own Generations
- Temporal Self-Rewarding Language Models: Decoupling Chosen-Rejected via Past-Future
- Approaching Human-Level Forecasting with Language Models
- Nexus: An Agentic Framework for Time Series Forecasting
- Large Language Models Cannot Self-Correct Reasoning Yet
- Think Twice Before Trusting: Self-Detection for Large Language Models through Comprehensive Answer Reflection
- Beyond Accuracy: The Role of Calibration in Self-Improving Large Language Models
- Post-Training Large Language Models via Reinforcement Learning from Self-Feedback