References Improve LLM Alignment in Non-Verifiable Domains

Paper · arXiv 2602.16802 · Published February 18, 2026
Frontier AI Risk & RSI

While Reinforcement Learning with Verifiable Rewards (RLVR) has shown strong effectiveness in reasoning tasks, it cannot be directly applied to non-verifiable domains lacking ground-truth verifiers, such as LLM alignment. In this work, we investigate whether reference-guided LLM-evaluators can bridge this gap by serving as soft “verifiers”. First, we design evaluation protocols that enhance LLM-based evaluators for LLM alignment using reference outputs. Through comprehensive experiments, we show that a reference-guided approach substantially improves the accuracy of less capable LLM-judges using references from frontier models; stronger LLM-judges can also be enhanced by high-quality (i.e., human-written) references. Building on these improved judges, we demonstrate the utility of high-quality references in alignment tuning, where LLMs guided with references are used as judges to self-improve. We show that reference-guided self-improvement yields clear gains over both direct SFT on reference outputs and self-improvement with reference-free judges, achieving performance comparable to training with ArmoRM, a strong finetuned reward model. Specifically, our method achieves 73.1% and 58.7% on AlpacaEval and Arena-Hard with Llama-3- 8B-Instruct, and 70.0% and 74.1% with Qwen2.5-7B, corresponding to average absolute gains of +20.2 / +17.1 points over SFT distillation and +5.3 / +3.6 points over reference-free self-improvement on AlpacaEval / Arena-Hard. These results highlight the potential of using reference-guided LLM-evaluators to enable effective LLM post-training in non-verifiable domains.

Introduction. Recently, Reinforcement Learning (RL) from Verifiable Reward (RLVR) (Liu et al., 2024a; Lambert et al., 2025) has shown strong effectiveness in improving LLMs’ reasoning capabilities. However, RLVR cannot be directly applied to non-verifiable domains, such as alignment tuning (Ouyang et al., 2022; Bai et al., 2022), because it is non-trivial to design verifiable/reliable rewards for these tasks. Consequently, RL from Human Feedback (RLHF) (Stiennon et al., 2020; Ouyang et al., 2022) or AI Feedback (RLAIF) (Bai et al., 2022) remains the predominant paradigm for LLM post-training in these domains.

There is a key difference between RLVR and RLHF: in RLHF/RLAIF, reward models or LLMsas-Judges (Zheng et al., 2024; Li et al., 2023; 2024) provide supervision signals, which typically evaluate outputs in a reference-free manner. In contrast, RLVR uses automatic verifiers that check outputs against gold-standard solutions (Liu et al., 2024a). This motivates our central question: can reference-guided LLM-evaluators act as soft verifiers to support effective RL for LLM alignment without external human or AI supervision?

We argue this is important for two reasons. First, it extends the reference-grounded supervision advantage of RLVR to non-verifiable domains where rule-based verifiers are infeasible. Second, high-quality reference outputs may be available even when preference labels are costly or unavailable, making direct reference-based improvement practically valuable.

To address this question, we develop LLM-judges1 that can effectively leverage reference outputs to provide supervision signals for preference optimization algorithms such as DPO (Rafailov et al., 2023). Critically, these reference-guided LLM-judges are used in a self-improvement manner, where an LLM serves as the judge to supervise its own training process (Yuan et al., 2024; Wu et al., 2024), so no external human or AI feedback is required. Recent work on rubrics-as-rewards (Gunjal et al., 2025; Huang et al., 2025) also leverages reference information (e.g., reference answers as supervision proxies or reference-based rubric construction), but its focus is rubric/reward design for RL. Our focus is different: we directly use external reference outputs to guide LLM-as-a-Judge decisions and then use these reference-guided judges for self-improvement training in alignment. Along this direction, Chang et al. (2025) has proposed to explicitly use reference outputs in RL for alignment tuning, however, its approach relies on using canonical metrics like BLEU (Papineni et al., 2002) instead of LLMs as evaluators.

To use LLM-judges for self-improvement training, we first need to ensure those judges are actually robust and accurate. Therefore, in §3, we first develop effective reference-guided LLM-judges for alignment evaluation. Several recent studies have explored guiding LLM-judges using references (Zeng et al., 2024; Lyu et al., 2024; Zhang et al., 2025; Krumdick et al., 2025). However, their evaluation settings are limited in the types of tasks considered and the number of LLMs used as judges, and a more systematic and comprehensive investigation is lacking (further discussed in §2). To this end, we first introduce targeted prompting strategies designed to leverage strong references for alignment evaluations. We then conduct comprehensive evaluations for the developed reference-guided evaluation method based on the prompting strategies. Notably, we find that naive incorporation of references without explicit guidance on how to use them yields only modest improvements, highlighting the importance of carefully designed prompting strategies. Specifically, our proposed method achieves a 6.8% absolute improvement over the reference-free baseline evaluated across 11 LLM-judges using reference outputs generated by a stronger LLM, GPT-4o (Hurst et al., 2024). Furthermore, frontier LLMs like GPT-4o can also be enhanced as judges when provided with high-quality human references.

Having developed the LLM-judges that can effectively leverage references, we apply them in a self-improvement setting for alignment tuning (§4). Specifically, we use the instructions in the widely used UltraFeedback dataset (Cui et al., 2023) to fine-tune Llama-3-8B-Instruct (Meta AI, 2024) and Qwen2.5-7B (Yang et al., 2024), with high-quality references generated by DeepSeek-V3 (Liu et al., 2024a). We conduct a reference-focused training process involving two stages – (1) first performing distillation, i.e., supervised fine-tuning (SFT) on the reference outputs, (2) then further improving the LLMs using DPO with the LLMs themselves as reference-guided LLM-judges to provide supervision.

Related work. LLM-as-a-Judge. Using powerful LLMs as automated evaluators (LLM-as-a-Judge) is a growing practice for scalable evaluation, especially in instruction-following tasks (Zheng et al., 2024; Li et al., 2023; Dubois et al., 2024). Benchmarks like MT-Bench and Arena-Hard (Zheng et al., 2024; Li et al., 2024) utilize strong LLMs (e.g., GPT-4) as judges, and this paradigm also supports training data annotation for preference optimization algorithms like DPO (Yuan et al., 2024; Rafailov et al., 2023). However, LLM judges are known to exhibit limitations such as positional and verbosity biases (Zheng et al., 2024; Zhu et al., 2023; Ye et al., 2024). Mitigation efforts include Chain-of- Thought (CoT) prompting (Wei et al., 2022), answer swapping (Shi et al., 2024), and developing more robust evaluation protocols (Zeng et al., 2024; Liu et al., 2024c). Our work builds on this by investigating reference-guided prompting to enhance LLM judge accuracy and robustness. Trivedi et al. (2024) explores improving judges via internal self-rationalization (Chain-of-Thought). Our work is orthogonal, focusing on external grounding via references to address knowledge gaps rather than reasoning gaps.

The Role of References in LLM Evaluation. Traditional NLG evaluation often relies on reference outputs (e.g., BLEU (Papineni et al., 2002), ROUGE (Lin, 2004)), but their role in LLM-as-a- Judge for alignment evaluation, where single ground-truth references are often insufficient, has been less explored. Recent work has begun revisiting references: LLMBar (Zeng et al., 2024) used prompts that guide LLM-judges to generate reference outputs before evaluation; HREF (Lyu et al., 2024) incorporated human-written responses and reported improved performance over reference-free methods. However, these studies are limited in both the number of LLMs evaluated and dataset scale.

Self-Improving LMs and Generative RMs. LLM-judges have also been used in model training, particularly in self-improvement settings where an LLM supervises its own training (Yuan et al., 2024; Wu et al., 2024; Yasunaga et al., 2024). A related line of work explores Generative Reward Models (Zhang et al., 2024; Mahan et al., 2024), where LLMs serve as reward models in preference optimization. Recent studies show that general-purpose frontier LLMs can perform competitively with finetuned discriminative reward models in this setting (Zhou et al., 2025; Frick et al., 2025).

Method. To enable effective use of references in improving LLM alignment evaluation, we first develop robust reference-guided evaluation methods for LLM-judges, and conduct comprehensive evaluations of them against strong baselines.

LLM-judges for alignment evaluation typically perform pointwise scoring of a single output given an instruction (Zheng et al., 2024), or pairwise comparison of two outputs (Li et al., 2024). In this study, we focus on the pairwise comparison setting, as it matches the annotation format of various high-quality human-labeled alignment datasets such as LLMBar (Zeng et al., 2024), and is directly applicable to preference optimization algorithms like DPO. While our primary analysis focuses on this pairwise setting, we also conducted experiments on pointwise scoring to ensure the robustness of our findings. We present these results in Appendix B, which confirms that reference-guided evaluation also improves performance in a pointwise scoring setting. To evaluate an LLM-judge, human annotations are typically used as ground truth. Specifically, in the pairwise comparison task, the LLM-judge’s evaluation accuracy is measured by the proportion of instances where the LLM-judge selects the same preferred output as the human annotators.

Existing reference-based prompting methods for LLM-judges (Zeng et al., 2024; Lyu et al., 2024) typically incorporate a reference output into the prompt, but provide limited guidance on how the judge should use it, resulting in only modest improvements over reference-free evaluation (as we show in Section 3.4). We show that more explicit instructions on reference utilization are needed, and introduce targeted prompting strategies to this end. We introduce targeted prompting strategies designed to effectively leverage reference answers in the LLM-as-a-Judge paradigm. As a baseline, we first introduce a strong reference-free prompting method, which we refer to as Ref-Free (Ours). This prompt is designed for direct pairwise comparison without relying on any external reference answer (its template is in Figure 9). Its general structure follows the base prompt proposed in Zeng et al. (2024). However, we design the prompt to specifically instruct the model to assess instructionfollowing quality along with other critical aspects such as factuality and verbosity. As results show, this method outperforms many existing baselines.

Building on the reference-based approach, we extend it to a reference-guided setting by instructing the LLM to assess which candidate output more closely aligns with the quality and content exemplified by the reference, while still addressing the original instruction. We refer to this method as RefEval. While prior work has proposed similar reference-guided prompting methods (Zeng et al., 2024; Lyu et al., 2024), our approach offers more explicit guidance on how the reference output should be used (See prompt design in Appendix A.2).

We show a snapshot of our core prompt in Figure 2, while the full prompt template is provided in Figure 10 in Appendix H. As shown in §3, this emphasis on reference utilization leads to clear improvements over previous methods.

To further emphasize the role of references, we design an additional prompting method, RefMatch (Figure 11), which instructs the LLM-judge to act primarily as a semantic and stylistic matcher, determining which candidate output more closely resembles the reference. Specifically, the LLM is explicitly instructed with: “Your goal is to determine which output demonstrates closer similarity to the reference.”

Having demonstrated the benefits of references in aiding LLM-judges’ evaluations, we now explore their utility in model training. Specifically, we consider a self-improvement setting where an LLM supervises its own training using preference optimization algorithms (Yuan et al., 2024; Wu et al., 2024). Unlike prior work, however, the LLM is provided with high-quality reference outputs to guide its evaluations, making this setup more practical.

Discussion. Table 3 highlights the following findings:

(1) SFT training on high-quality reference outputs is more effective than directly running preference optimization from the base model with a finetuned reward model. Specifically, DSV3-Distill generally outperforms ArmoRM-Base and is competitive on Arena-Hard for Qwen2.5-7B-SFT. This underscores the benefits of strong reference outputs.

(2) LLMs can effectively serve as their own judges for preference optimization (i.e., selfimprove), which is demonstrated by the substantial improvement of RefFree over DSV3- Distill. On Llama-3-8B-Instruct, this gain is +13.6 on AlpacaEval and +11.6 on Arena-Hard; on Qwen2.5-7B-SFT, it is +16.3 and +15.3, respectively.

(3) References help LLMs better self-improve, as RefEval, the model trained with the referenceguided self-judge, consistently outperforms Ref- Free. It also performs much more strongly than the traditional reference-based metrics, ROUGE and BERTScore, and achieves comparable or better performance compared to the finetuned reward model ArmoRM. As shown in Table 4, reference-guided self-improvement yields additional gains over RefFree on both models and both benchmarks. The largest gains over DSV3-Distill reach +21.2 on AlpacaEval and +17.6 on Arena-Hard. It shows that the LLM-judges’ improvement from references observed in §3 can result in substantial improvement in training. demonstrate a clear advantage of our trained models, validating the effectiveness of leveraging high-quality references.

Understanding the Benefits of Reference-Based Self-Improvement. To better understand the distinction between reference-based and reference-free supervision, we use GPT-4o to categorize the instructions in AlpacaEval and Arena-Hard into four types: Coding&Math, Creative Tasks, Information Seeking, and Reasoning&Planning.

We then compare the performance of the RefFree and RefEval models across each category.6 Figure 3 shows that for both Llama-3-8B-Instruct and Qwen2.5-7B-SFT, reference-based supervision yields a substantial improvement in the Coding&Math category. However, its benefit on Creative Tasks is less significant for Qwen2.5-7B-SFT, while remaining considerable for Llama-3-8B-Instruct. We posit that this is because leveraging references effectively in open-ended tasks is more challenging, and doing so requires more extensive post-training (as in Llama-3-8B-Instruct) rather than the standard SFT (as in Qwen2.5-7B-SFT).

Conclusion. In this study, we investigate whether high-quality reference outputs can enable effective LLM alignment tuning. Across five datasets, we show that high-quality references can consistently improve LLM-judge performance. Using the developed reference-guided LLM-judges in alignment tuning, we demonstrate that they can lead to effective semi-self-improvement by using high-quality references, even achieving performance comparable to that of trained reward models. Our findings highlight the potential of leveraging references to improve LLMs in non-verifiable domains, while reducing the methodology gap between RLHF/RLAIF and RLVR for LLM post-training. Along this direction, we believe future work should focus on exploring the effectiveness of references in more specialized domains that require domain expertise and knowledge, and on developing reward models that specifically utilize references in such scenarios.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Does reinforcement learning create genuinely new reasoning capabilities or only refine existing ones? How do reward signal properties affect model reasoning and safety? What makes reasoning traces effective supervision even when they're incorrect? How can we reduce inherent biases in LLM-based evaluation judges? How can evaluations be made robust against model reward hacking? Why does polished AI output gain credibility despite fundamental verifiability problems? Can external verification systems adequately replace learned reasoning in AI outputs? When should retrieval systems decide to fetch new information? Do language models encode knowledge that influences generation, or primarily imitate surface patterns? Why does AI verification capability persistently exceed generation capability? Can minimal training unlock latent reasoning already present in base models? What prevents LLMs from applying their reasoning knowledge to improve outputs? Why do multi-agent systems reach premature consensus without genuine deliberation? How do curriculum design and feedback approaches affect model learning? How should retrieval strategies adapt to multi-step reasoning demands? What limits recursive self-improvement in autonomous AI systems?