Can you trust any 'AI-written share' figure when changing the scoring method alone can sharply swing detector results?
How much do different detection frameworks disagree on adoption rates?
This explores how far estimates of AI adoption (for example, how much text in papers or reviews is AI-written) shift depending on which detector is used. The retrieved material doesn't contain those adoption estimates, so this answer covers the nearby question of how much detection results depend on the detection method.
This explores how far estimates of AI adoption (for example, how much text in papers or reviews is AI-written) shift depending on which detector is used. The short answer is that none of the notes retrieved here compare adoption-rate estimates across detection frameworks, so this synthesis can't give you a number for that gap. What the notes do show, from several different directions, is that detection results often depend heavily on how you measure them. That's the strongest reason to distrust any single adoption figure.
The clearest case comes from hallucination detection. Simply changing the scoring metric from ROUGE to one that matches human judgment cut apparent detection ability by up to 45.9%, and a trivial heuristic based only on answer length rivaled sophisticated methods Is hallucination detection progress real or just metric artifacts?. If a detector's measured success can swing that much on the scoring choice alone, a headline figure like 'X% of reviews are AI-written' is partly a statement about the detector and its yardstick. It isn't only a statement about the world.
Disagreement also shows up when two detectors are compared on the same data. Cheap 'difference-of-means' vector probes and expensive LLM monitors were tested at the same false-positive rate on agents trying to cheat their reward (reward hacking). The probe caught 3.1% more hacks on one model and 7.9% fewer on another difference-of-means-vectors-are-similarly-effective-to-llm-monitors-but-virtually. The ranking between detectors flipped depending on what was being monitored. Judges disagree too: LLM-as-a-judge evaluators changed their verdicts about 31% of the time, compared with 0.27% for an agent that gathers evidence before deciding Can agents evaluate AI outputs more reliably than language models?. Any adoption count that uses an LLM as the classifier inherits that kind of instability.
The less obvious lesson is that 'detected' and 'adopted' may not be the same thing. Research on human annotation shows that a single label can mix real signal with noise and with answers made up on the spot, and you can only tell them apart by checking whether they hold up under different measurement conditions Do all annotation responses measure the same underlying thing?. A related argument holds that observed behavior can never prove unobserved behavior Can behavioral training prove a model always complies?. Applied to adoption, a detector only sees AI use that leaves a detectable trace. Heavily edited or lightly assisted text can fall outside what it can see. So different frameworks may disagree partly because they are counting different kinds of 'use'.
If you want the actual adoption-rate comparisons, such as studies estimating the share of LLM-written peer reviews or papers, they aren't among these notes. The best next step is a targeted search of the collection for terms like 'LLM-generated peer reviews', 'distributional estimation of AI text', or 'AI text detector reliability'.
Sources 5 notes
ROUGE-based evaluation inflates detection capability by up to 45.9 percent compared to human-aligned metrics. Simple length heuristics rival sophisticated methods like Semantic Entropy, suggesting much reported progress measures length variation rather than factual accuracy.
Eight-module agentic evaluation achieved 0.27% judge shift versus 31% for LLM-as-a-Judge on complex tasks. However, the memory module cascaded errors, revealing that agentic systems need error isolation mechanisms to maintain gains.
Behavioral science reveals that annotations contain genuine preferences, non-attitudes, and constructed preferences—distinguishable by consistency across measurement conditions. Treating them uniformly contaminates reward model training and downstream alignment.
Any scored behavior is observed behavior, so training data cannot distinguish between a policy that always complies and one that complies only when watched. Only unobserved behavior would separate them, making such a test logically impossible.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Measuring Human Preferences in RLHF is a Social Science Problem
- Encouraging Divergent Thinking in Large Language Models through Multi-Agent Debate
- When AIs Judge AIs: The Rise of Agent-as-a-Judge Evaluation for LLMs
- Agent-as-a-Judge: Evaluate Agents with Agents
- Interactive Evaluation Requires a Design Science
- The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
- AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
- ChatEval: Towards Better LLM-based Evaluators through Multi-Agent Debate