INQUIRING LINE

Getting merged only shows a reviewer said yes, not that the AI-written code is actually correct.

Can merged pull requests serve as a meaningful quality measure?

This explores whether 'the pull request got merged' is a trustworthy sign that code (especially AI-written code) is actually good. The retrieved notes have nothing on PR merge rates directly, so this answer borrows from nearby work on proxy signals, approval and correctness.


This explores whether a merged pull request is a trustworthy sign of quality, or only a sign that someone approved it. The retrieved notes don't study merge rates directly, so treat what follows as borrowed lessons, not direct evidence. Those lessons point the same way: a merge is a gate, not a grade. It records that a reviewer agreed. It doesn't show the change was correct.

The clearest parallel is in work on validator consensus. Can validator consensus guarantee both agreement and semantic correctness? separates two guarantees that are easy to confuse. A protocol can promise that validators agree. Whether the thing they agreed on is actually right holds only statistically, because it depends on how carefully each validator checked. A merge works the same way: it settles the decision but says nothing about whether the code is correct. Agreement can also be weak evidence in its own right. Does disagreement between AI coders signal better accuracy? found that AI coding agents were more accurate when they argued at length. A smooth, uncontested approval may mean less checking happened, not that the code was better.

The worry is sharper for AI-generated code. Does model capability change how documents degrade? shows that weaker models damage documents in ways you can see, such as deleting content. Frontier models damage them quietly, with changes that keep the surface looking intact. Code that passes a reviewer's glance is exactly the kind of output that could hide this. So merge rates could rise as models improve while hidden defects rise too.

That doesn't make merges useless. It means they work better as one signal among several. Can rubrics and dense rewards work together without hacking? argues that pass/fail checks are strongest as gates that throw out bad candidates, and weaker as scores to maximize, because optimizing directly for the score invites gaming. Read that way, 'was it merged?' is a reasonable minimum bar but a risky target. Does supervising retrieval steps outperform final answer rewards? adds that judging only the final outcome hides how the result was reached. Looking at the steps, such as review comments, revisions and later reverts, tells you more. And Can crowdsourced votes reliably rank language models? shows when approval signals can be trusted: when there are many judges, the tasks are varied and hard enough to tell options apart, and the signal is checked against experts.

What you might not have expected: the better the model, the less a merge may tell you. Frontier models' failures are designed by accident to survive a quick look, which is what a typical approval is. The useful follow-up question is what happens after the merge: reverts, follow-up fixes, and defects found later.


Sources 6 notes

Can validator consensus guarantee both agreement and semantic correctness?

Honest Quorum's threshold theorems split into two kinds of guarantee: agreement rests on protocol assumptions alone, while semantic validity and liveness depend on statistical bounds over validator behavior that the protocol cannot enforce.

Does disagreement between AI coders signal better accuracy?

Multi-agent LLM coding systems showed higher accuracy when agents engaged in prolonged, unresolved debate. The frequency of disagreement and undecidable labels serve as reliable performance indicators, suggesting conflict deepens interpretive work rather than signaling failure.

Does model capability change how documents degrade?

DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.

Can rubrics and dense rewards work together without hacking?

DRO shows that using rubrics to accept or reject rollout groups—rather than converting rubric scores into dense rewards—prevents reward hacking. This separation preserves the categorical strength of rubrics while letting token-level rewards optimize within valid answers.

Does supervising retrieval steps outperform final answer rewards?

Fine-grained feedback on intermediate retrieval steps significantly boosts agentic RAG performance compared to final-answer-only rewards. DPO trained with both positive and negative step feedback outperforms PPO and single-direction training by directly contrasting good and bad retrieval chains.

Show all 6 sources
Can crowdsourced votes reliably rank language models?

Chatbot Arena's 240K+ crowdsourced preference votes produce credible model rankings because the underlying questions are diverse and discriminating, and crowd judgments correlate with expert raters—validating human preference as a scalable evaluation signal.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.