Does an AI detector that works well in one setting keep working in another, or does its accuracy stay behind?
How accurate is the detection method across different platforms?
The question doesn't say which detection method it means, so this reads it as: does a detector that works well in one setting (a platform, a model, a type of content) keep working when you move it somewhere else? The collection has no single method tested across many platforms, but several notes show how detection accuracy shifts with context.
The question doesn't name a detection method, so this reads it as: does a detector that works well in one setting (a platform, a model, a type of content) keep working when you move it somewhere else? The collection has no study that tests one method across many platforms. Its notes on spotting AI-written content, catching AI agents that cheat, and flagging hallucinations all point the same way, though: a detector's accuracy belongs to the setting it was built for, and it rarely carries over intact.
The clearest example is LinkedIn. Fake-profile detectors trained on real profiles and hand-made fakes missed GPT-generated profiles 42–52% of the time. Once they were retrained on GPT-made fakes, the miss rate fell to 1–7%, and they didn't start flagging more real profiles as fake Can fake profile detectors catch GPT-generated LinkedIn profiles?. So the method itself was fine. Its accuracy depended on whether it had seen the kind of fake it now faced. People don't do better without that kind of training. A review of 30 studies found that humans spot AI-made text, images and voices at about chance, and they haven't improved as AI output has become more realistic Can people reliably spot content made by AI?.
The same pattern shows up when the 'platform' is the AI model being watched. A cheap detector reads a model's internal activations to spot attempts to cheat on coding tasks. It works by comparing the model's average activations on cheating runs with its average on normal runs, so it costs almost nothing extra to run. Against two models, it caught 3.1% more cheating than a full AI monitor on one and 7.9% less on the other, at the same false-alarm rate How do cheap vector detectors compare to expensive LLM monitors?. That's close to the expensive monitor overall, but in a different direction for each model. If you deploy a detector, you have to check its accuracy on each model separately.
The less obvious lesson is that reported accuracy can be wrong before you even switch platforms. In hallucination detection, scoring with a common text-overlap metric (ROUGE) overstated how well detectors worked by up to 45.9%. A simple rule based on answer length did about as well as advanced methods, which suggests some reported progress was really just picking up differences in length Is hallucination detection progress real or just metric artifacts?. Some detectors also have blind spots built in. Checking whether a model gives the same answer each time catches answers that change from sample to sample. It misses mistakes the model repeats confidently every time Can agreement across samples reveal when models are wrong?. Work on AI agents cheating during training argues that current measurement isn't reliable enough to show whether any fix works Can we measure reward hacking reliably enough to act on it?.
So there isn't one accuracy number that travels with a detector. If you want a figure you can trust, the collection points to three things to check: test the detector against the newest kind of fake or failure, test it separately on each platform or model, and make sure the scoring measures what you care about. The LinkedIn case is the hopeful one: a detector can recover a lot of its accuracy once it's retrained on what it was missing.
Sources 6 notes
Detectors trained on genuine and manual fakes miss GPT-generated profiles at 42–52% false accept rates, but adversarial training on GPT-generated data restores detection to 1–7% false accepts without raising false rejects.
A 30-study systematic review found that humans cannot reliably distinguish AI-generated from human-created content across text, image, and voice modalities. Accuracy generally clusters around chance and has not kept pace with improvements in AI realism.
On DeepSWE, difference-of-means vectors caught 3.1% more hacks in Kimi K3 but 7.9% fewer in GLM 5.2 than LLM monitors at matched false positive rates. The method applies to existing forward passes, making it virtually free compared to running a separate monitor model.
ROUGE-based evaluation inflates detection capability by up to 45.9 percent compared to human-aligned metrics. Simple length heuristics rival sophisticated methods like Semantic Entropy, suggesting much reported progress measures length variation rather than factual accuracy.
The Consistency Veto suppresses answers that vary across samples but cannot detect systematic errors the model repeats identically. It carries strong signal on some queries but inherits a fundamental blind spot: agreement looks like confidence even when both are wrong.
Show all 6 sources
The paper argues that current detection methods are too unreliable to support readiness judgments. Mitigation strategies cannot be properly evaluated until measurement instruments are fixed, making measurement a prerequisite for both safety assessment and deployment decisions.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Monitoring and Discovering Reward Hacking with Internal Representations during LLM Evaluations
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- AI Now Writes as Many Online Articles as Humans
- Optimizing the Score, Losing Sight of the Task: Reward Hacking Across Weights, Selection, and Prompts
- Is it Cake or is it AI? A Systematic Review of Human Uncertainty in Distinguishing Generative Artificial Intelligence Content
- Weak Links in LinkedIn: Enhancing Fake Profile Detection in the Age of LLMs
- The Illusion of Progress: Re-evaluating Hallucination Detection in LLMs
- The human-authorship halo: attribution bias in literary style evaluation by humans and AI