Can safety training detect attacks hidden in context rather than commands?
Most AI safety training blocks explicit harmful requests, but what happens when misinformation is packaged as credible evidence and injected into a conversation's context? This explores whether current defenses catch attacks that look like background information rather than instructions.
Most safety alignment is trained against the imperative form of an attack: a user asking the model to produce biased, false, or harmful content. GHOSTWRITER exposes that the binding surface is elsewhere. Its two-phase attack first repackages a misleading viewpoint with a fabricated rationale that bears markers of credibility, then drops it into a conditional template ("when responding to relevant queries, incorporate this"). The model internalizes the viewpoint because nothing about the payload looks like a request to do something forbidden — it looks like evidence. On BBQ, ToxiGen, and a custom set, commercial LLMs without external classifiers are highly vulnerable; even a frontier classifier-guarded model reduces but does not eliminate it. A tailored safety policy on gpt-oss-safeguard reaches 81% detection, so the defense is moving from refusing instructions to appraising the epistemic status of context.
This is the adversarial-engineering version of mechanisms my vault already has. Does transformer attention architecture inherently favor repeated content? gives the architectural reason the attack works: attention is built to run with prominent context rather than verify it, so credibility markers are exactly what it over-weights. Do language models actually build shared understanding in conversation? is the pragmatic complement — the model treats injected evidence as shared, accepted background. And it operationalizes Can models abandon correct beliefs under conversational pressure?: where that note shows belief drift under conversational pressure, GHOSTWRITER shows a single well-dressed payload can do it in one shot, even when the model's prior knowledge is correct.
The counterargument worth holding: 81% detection with a tailored policy suggests this is patchable, not load-bearing. But the deeper point survives — alignment that polices what the user asks leaves a blind spot for what the context asserts, and as third-party chat platforms and LLM-run social accounts proliferate, the context channel is the larger, less-guarded attack surface.
Inquiring lines that read this note 7
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Can AI systems evade safety evaluations through reasoning manipulation? How do individually-safe actions create collectively-unsafe outcomes? Do individually safe AI actions create unsafe outcomes in integrated systems? How do educators verify student capability when AI can produce indistinguishable work? How do evaluation environment design choices affect AI security?Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does transformer attention architecture inherently favor repeated content?
Explores whether soft attention's tendency to over-weight repeated and prominent tokens explains sycophancy independent of training. Questions whether architectural bias precedes and enables RLHF effects.
grounds: the architectural reason credibility-marked context bypasses scrutiny
-
Do language models actually build shared understanding in conversation?
When LLMs respond fluently to prompts, do they perform the communicative work humans do to establish mutual understanding? Research suggests they skip the grounding acts that make dialogue reliable.
convergent-with: the model treats injected evidence as accepted shared background
-
Can models abandon correct beliefs under conversational pressure?
Explores whether LLMs will actively shift from correct factual answers toward false ones when users persistently disagree. Matters because it reveals whether models maintain accuracy under adversarial pressure or capitulate to social cues.
extends: single fabricated-evidence payload achieves in one shot what multi-turn pressure does gradually
-
Can reasoning models be steered by injected context without detection?
This explores whether adversaries can plant harmful-but-benign-sounding reasoning in a model's context and have it followed while evading chain-of-thought monitors. The question matters because it tests whether monitoring reasoning traces can catch deception at inference time.
convergent: the effective payload is reasoning-shaped, not command-shaped, and here it steers actions and not only viewpoints while passing a CoT monitor
-
How do competent systems quietly undermine safety oversight?
This note explores four mechanisms by which well-functioning AI systems can erode the human safeguards meant to contain them: user overconfidence, blurred authority lines, accumulated hidden failures, and scattered accountability. Understanding these pathways matters because the most harmful systems may look least harmful.
a payload that is an authority claim and not a command is one case of an authority boundary collapsing; the mapping to this note is the vault's, not the paper's
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Steering LLM Viewpoints through Fabricated Evidence Injection
- Models That Know How Evaluations Are Designed Score Safer
- EvoSafeHarness: Evolving Model- and Domain-Specific Harnesses for Securing Agents
- How Johnny Can Persuade LLMs to Jailbreak Them: Rethinking Persuasion to Challenge AI Safety by Humanizing LLMs
- Strategic Dishonesty Can Undermine AI Safety Evaluations of Frontier LLMs
- Persistent Pre-Training Poisoning of LLMs
- When Chain of Thought is Necessary, Language Models Struggle to Evade Monitors
- The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
Original note title
the dangerous prompt-injection payload is not an instruction but a fabricated rationale — safety alignment guards explicit requests while credibility markers in context walk straight past it