Fake-profile detectors may only catch the kinds of fakes they were trained on, and nobody has tested profiles a person and AI co-wrote.
Can platforms trust these detection rates on profiles created with human-LLM collaboration?
This explores whether the strong fake-profile detection results (detectors catching GPT-written LinkedIn profiles once retrained on them) still hold when a profile is neither fully human nor fully machine-written, but co-written by a person editing or prompting an LLM.
This explores whether the strong fake-profile detection results still hold for hybrid profiles, where a person drafts with an LLM, edits its output, or mixes their own writing with generated text. The short answer: the corpus has no study that tests hybrid profiles directly, so nobody has measured these rates for co-written profiles. What the corpus does show is a pattern that should make platforms cautious about carrying the numbers over.
The key finding is that detectors catch only the kinds of fakes they were trained on. In Can fake profile detectors catch GPT-generated LinkedIn profiles?, detectors trained on real profiles and hand-made fakes let 42–52% of GPT-generated fakes through. Retraining on GPT fakes cut that to 1–7% without wrongly flagging more real users. That is an encouraging result, but it is really a result about matching the training data to the attack. A profile that a human has rewritten from an LLM draft (or an LLM has polished from a human draft) is a third kind of text, and it was in neither training set. The study's own story, where the detector fails until it has seen the new type of fake, is the best reason to expect hybrids to open a new gap. Until someone tests them, the 1–7% figure describes pure GPT fakes only.
A second problem comes from looking at how AI systems judge people when there isn't much signal to go on. When web-browsing LLMs guessed demographics from social media accounts, they did well on active accounts. On sparse ones they fell back on stereotypes, and their errors were skewed by gender and politics Can LLMs predict demographics from social media usernames alone?. Separately, every one of 12 tested LLMs made claims about users that the evidence didn't support, 35–49% of the time. The models that rated themselves as most careful were actually the worst offenders Do large language models fabricate user attributes beyond available evidence?. If platforms use LLM-based judges to decide whether a borderline hybrid profile is real, those judges are likely to settle ambiguous cases with confident guesses, and those guesses may fall harder on some groups of users than others.
The surprising angle is that "AI-assisted" may be the wrong thing to detect in the first place. In Do LLM raters show hidden demographic preferences that disclosure erases?, simply disclosing that AI was involved changed how both LLM and human raters scored the same writing. Human raters penalized disclosed AI use across the board, and the LLM raters' hidden demographic preferences disappeared. Platforms face the same tension. Huge numbers of real people now write their profiles with LLM help. A detector that flags "LLM-flavored text" will increasingly catch honest users, while a detector tuned to ignore AI help may miss fakes built the same way. For platforms, "Is this profile fake?" and "Did an LLM help write this?" are separate questions, and they are coming apart.
In short, platforms can trust these rates for the threat they were measured on: profiles that GPT generated from start to finish. For co-written profiles, the corpus suggests the rates will go wrong in predictable ways. Expect new misses at first, overconfident judgments on thin profiles, and a growing risk of flagging real people who used AI for help. A sounder approach is to detect fake identities through signals other than the writing style, and to keep retraining on new kinds of fakes as they appear.
Sources 4 notes
Detectors trained on genuine and manual fakes miss GPT-generated profiles at 42–52% false accept rates, but adversarial training on GPT-generated data restores detection to 1–7% false accepts without raising false rejects.
Evaluated on 1,384 survey participants and 48 synthetic accounts, web-browsing LLMs successfully predicted gender, age, and political orientation from X usernames and profiles alone. The models showed systematic gender and political biases specifically against low-activity accounts, relying on stereotype-driven defaults when content was sparse.
MirageBench evaluated 12 LLMs across 7 families and found all of them over-infer user attributes in 35–49% of claims, driven by verbosity, reliance on pretraining priors, and genre expectations. Models that self-assess as over-inferring less actually over-infer more when judged independently.
GPT-4o-mini showed pronounced preference for Black authors and Qwen2.5-7B-Instruct favored women authors when AI use was undisclosed, but both preferences vanished under disclosure. Human raters showed uniform disclosure penalties regardless of author demographics.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads
- LLM or Human? Perceptions of Trust and Information Quality in Research Summaries
- LLM Targeted Underperformance Disproportionately Impacts Vulnerable Users
- Monitoring AI-Modified Content at Scale: A Case Study on the Impact of ChatGPT on AI Conference Peer Reviews
- Web-Browsing LLMs Can Access Social Media Profiles and Infer User Demographics
- Weak Links in LinkedIn: Enhancing Fake Profile Detection in the Age of LLMs
- Penalizing Transparency? How AI Disclosure and Author Demographics Shape Human and AI Judgments About Writing
- Understanding the Role of User Profile in the Personalization of Large Language Models