SYNTHESIS NOTE
Topics›Expertise in the Age of AI Content›this note

Why do text embeddings fail faster under LLM attack?

Text-only profile embeddings collapse under LLM adversarial attacks while numerical features remain stable. Understanding this gap could reveal whether the fragility stems from how text encodes meaning or from something specific to how LLMs generate profiles.

Synthesis note · 2026-10-06 · sourced from Expertise in the Age of AI Content

The paper's ablation finds that feature type changes how a LinkedIn fake profile detector fails. Detectors built on text-only embeddings are the most fragile under LLM attack, detectors built on numerical profile features are sturdier, and detectors that fuse the two are the most robust. The introduction puts it as "text embeddings alone are fragile under LLM attack, whereas numerical profile features remain sturdier; their fusion yields the most robust detector." The abstract ranks the three configurations in the same order: combined numerical and textual embeddings first, numerical-only second, textual-only last. The paper also presents the numerical side as the cheaper option: the "compact 17-dimensional numerical representation offers a lightweight alternative for detection in resource-constrained settings."

The adversarial training results are asymmetric by feature type. Text-only models trained with GPT-4-assisted data cut false accept rates to 2.59% to 8.33% on GPT-3.5 profiles and 3.12% to 8.33% on GPT-4 profiles. Numerical models gained most on the generator they were trained against. GPT-3.5-assisted training raised F1 on GPT-3.5 profiles from 78.67%-79.21% to 82.55%-84.08%, and GPT-4-assisted training reached F1 up to 96.75% on GPT-4 profiles. On the full 167-dimensional set, GPT-4-assisted training helped GPT-4 profiles substantially but had a limited effect on GPT-3.5 profiles; the example given is Flair with XGBoost, whose false accept rate fell only from 38.52% to 36.11%. The Discussion's explanation is that "textual features tend to generalize more effectively across model variants and adversarial scenarios, whereas numerical features seem to encode generation-specific artifacts."

This qualifies the first note. Can fake profile detectors catch GPT-generated LinkedIn profiles? treats adversarial training as one fix, and this note says which feature sets carry it and how far each generalizes. It also shares a warning with Why do fake news detectors flag AI-generated truthful content?, which locates detector bias in LLM linguistic patterns. Both point to the surface text channel as a weak point for detection. The two excerpts test different things, though, so the link is a shared warning, not a shared finding.

The excerpt does not establish why numerical features hold up better. "Seem to encode generation-specific artifacts" is the authors' interpretation and is not tested here. The excerpt does not identify which of the 167 numerical features carry the signal, and it does not show how the same features behave against non-GPT generators. It also leaves a tension unresolved. Text-only embeddings are the weakest configuration under attack, yet after adversarial training they generalize best across GPT variants. The implication is that feature choice depends on the threat. For a detector facing unseen generators, the excerpt supports fused features within the GPT family it tested, and nothing beyond that.

Inquiring lines that read this note 2

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

How reliably can humans and AI detectors identify machine-generated text? Can persona profiles improve LLM prediction accuracy and consistency?

Related concepts in this collection 2

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 75 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

text-only profile embeddings degrade sharply under LLM attack while numerical features hold and fusing both gives the most robust LinkedIn detector