Contaminating just 0.1% of an AI's training data can plant behaviors that still survive its later safety training.
How much does training data composition shape security model performance?
This explores how what goes into a model's training data (how much of it, from where, and how it's mixed) affects a model's security: both how well it resists attack and how well it does security-related work. The corpus speaks much more to the first than the second.
This explores how what goes into a model's training data affects its security: how easily the model can be compromised, and how well it handles security-related tasks. The corpus has little direct evidence on the second part. It has no studies of models trained to find vulnerabilities or classify threats where the data mix is varied and performance is measured. What it does have is a sharper version of the question: training data composition is itself an attack surface, and small parts of the mix can matter far more than their size suggests.
The clearest example is that contaminating just 0.1% of pretraining data is enough to plant behaviors that survive later safety training How much poisoned training data survives safety alignment?. The surprising part is which behaviors survive. Attacks that make the model refuse service, leak its context, or shift its beliefs persisted through alignment. Jailbreaking was the one attack that alignment reliably removed. So the safety training most people picture as the defense is tuned to one kind of threat and largely misses the others. A related threat class goes after distribution rather than the data itself. Backdoored checkpoints or hijacked model hubs can quietly insert promotional or malicious content while leaving benchmark accuracy untouched Can language models be hijacked to embed hidden advertisements?. That means the usual way of checking a model, its test scores, can't show whether its data has been tampered with.
A second lesson comes from research that isn't framed as security at all: later training amplifies some parts of the pretraining mix and erases others. RL post-training tends to lock onto one dominant format from pretraining within the first epoch and suppress the alternatives. Which format wins depends on model scale, not on which one works best Does RL training collapse format diversity in pretrained models?. Similarly, only about 20% of tokens carry most of the learning signal during reasoning training Do high-entropy tokens drive reasoning model improvements?. Neither study looks at security directly. But if a small slice of data or tokens can steer a whole training run, then a small poisoned or skewed slice could have outsized effects too. This is an inference, not a tested result.
If the goal is to build a model that's good at a security domain, the domain-adaptation research warns that "more and better data" isn't the whole answer. Each adaptation method has domain-specific conditions where it works best, and visible gains often come with hidden losses in reasoning faithfulness or flexibility How do domain training techniques actually reshape model behavior?. Data that is objectively better can still hurt a smaller model if it sits beyond what that model can learn, so students do better when they filter a teacher's refinements by fit Does teacher-refined data always improve student model performance?. Including explicit wrong examples also helps. In one study, DPO on a teacher's correct and incorrect function calls taught small models better than imitation alone Can small models match large models on function calling?. Security work, with its strict formats and costly mistakes, looks like the kind of domain where that would matter.
In short, the corpus can't tell you how much data composition shifts a security model's benchmark scores. It does suggest a more useful question: which small parts of the mix are doing the steering, and would anyone notice if they changed? For the defensive side of this, the work on judging agent actions as linked coordination episodes rather than isolated events is a natural next stop Should defence units span multiple executions and agents?.
Sources 8 notes
Denial-of-service, context extraction, and belief manipulation attacks persist through standard safety alignment at 0.1% poisoning rates, while jailbreaking attacks are successfully suppressed, contradicting sleeper agent persistence hypotheses.
Research identifies Advertisement Embedding Attacks as a distinct threat class that injects promotional or malicious content via hijacked distribution platforms or backdoored checkpoints, leaving accuracy untouched while corrupting output integrity. The attack is economically motivated and self-inspection defenses can detect injected content without retraining.
Controlled experiments show RL consistently amplifies one format distribution from pretraining within the first epoch while collapsing alternatives. The winning format depends on model scale, not necessarily performance, and is largely hidden when starting from proprietary pretrained models.
Only ~20% of tokens exhibit high entropy as pivotal reasoning decision points; RLVR primarily adjusts these forking tokens. Training exclusively on them matches or exceeds full-gradient performance, revealing that the minority carries the learning signal.
Research shows every adaptation method—from parameter-efficient tuning to knowledge graph curricula—has optimal conditions tied to specific domains. The key finding: visible benefits like performance gains often come with hidden degradation in reasoning faithfulness, capability transfer, and format flexibility.
Show all 8 sources
Teacher-refined data degrades performance when it exceeds the student's learning frontier, even if objectively higher quality. Students should filter refinements using their own statistical profile to retain only compatible improvements.
Small models fine-tuned via DPO on correct and incorrect function-calling examples from a large teacher model achieve high accuracy on logical and mathematical tasks. DPO's explicit negative examples directly target the rigid output format failures where SFT alone underperforms.
The operational unit of defence should be a set of actions linked by observed transfers, task authority, and response history, with membership revised as evidence accumulates. Isolated review loses relevant context that spans multiple executions.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Echo Chamber: RL Post-training Amplifies Behaviors Learned in Pretraining
- Persistent Pre-Training Poisoning of LLMs
- AI4AI at Test-Time: Strong-to-Weak Capability Transfer via Harnesses
- Sharpening Tax in Post-Training
- AutoPrompt: Eliciting Knowledge from Language Models with Automatically Generated Prompts
- Anthropic Risk Report: August 2026 (redacted)
- Beyond the 80/20 Rule: High-Entropy Minority Tokens Drive Effective Reinforcement Learning for LLM Reasoning
- Attacking LLMs and AI Agents: Advertisement Embedding Attacks Against Large Language Models