Line of inquiry
Inquiring lines›How do we keep AI systems safe and…›How do adversarial attacks exploit…›this line of inquiry
Can AI systems evade safety evaluations through reasoning manipulation?
A broader line of inquiry — a family of 80 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 80
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Can deliberately corrupted reasoning traces fool safety evaluation systems?
- How can simple prompt injection attacks extract reasoning trace content?
- Does optimizing against CoT monitors inevitably produce obfuscated reasoning?
- Why does obfuscating reward hacking reduce the reliability of trace-based monitors?
- Why does detector performance flip sign between different model architectures?
- Can activation-level monitoring catch hacks that leave no clean trace?
- Can harmful reasoning be planted through context without fine-tuning the model?
- Does iterative optimization against CoT monitors reliably produce obfuscated evasion?
- How does the copyable-rule squeeze interact with the false-alert cost squeeze?
- Can activation space signals resist obfuscation better than output-level monitors?
- Do CoT-based monitoring systems degrade when models optimize against them?
- Can membership inference attacks reliably detect training data exposure?
- How do adversarial triggers bypass the protections of longer reasoning chains?
- What detection methods can catch each distinct CoT bypass strategy?
- How can model routing and provenance become an attack surface?
- Can minimal adversarial triggers disrupt reasoning across multiple unrelated queries?
- Does activation masking prevent the decoder from taking interpretability shortcuts?
- How do server-side filters hide their role in zero attack success?
- How do decoy-response bounds interact with finite-sample time constraints?
- What makes reasoning-shaped payloads more effective than command-shaped attack prompts?
- How much does training against monitors teach models to obfuscate?
- Can planted test cases remain unrecognizable to adaptive optimizers over time?
- Can activation decoders discover hidden system prompts from user-model conversations?
- Why do standard safety filters miss advertisement embedding attacks?
- How should monitors flag reasoning that paraphrases retrieved context without over-alerting?
- Can defenders detect attacks that probe scanner feedback as a learning signal?
- How do chain-level defenses differ from per-skill scanner detection approaches?
- Can verifiable execution traces replace fluent output as a training signal?
- Can re-scoring detect subliminal prompt injection without explicit semantic content?
- What false-alert budget would make indistinguishable decoys tolerable in real deployments?
- Why do paraphrasing defenses fail against subliminal prompt attacks?
- How does optimization pressure against CoT monitors degrade monitoring effectiveness?
- Do attackers adapt their plans when monitors deepen their reasoning budget?
- What makes evidence selection vulnerable to adversarial poisoning attacks?
- Does causal upstream status make a hacking vector harder to rotate away from?
- Do synthetic attack traces in papers reflect real adversary behavior?
- What role does a forged approval claim play compared to an explicit instruction?
- Can attack success rates hide server-side filtering or other non-adversarial defenses?
- Can consistency training defend against adversarial text injection attacks?
- How do backdoored open-source checkpoints enable covert advertising at scale?
- Can a deployed system verify the actual identity of the model that responded?
- Can entropy signatures alone detect whether context was model-generated or externally prefilled?
- How do interchangeable encrypted blocks enable cross-model attacks?
- Can phase-aware static taint analysis scale across different benchmark task types?
- What detection rate is needed to make evidence-injection attacks impractical at scale?
- Why do small training data contaminations persist through alignment for most attack types?
- Can knowledge poisoning attacks succeed with less than 0.05 percent modified text?
- What unnamed exploits do models discover in training environments?
- Do record-tampering and covert sabotage share a common underlying mechanism?
- Why should identifying the spy correlate with output quality?
- Why are expensive rankers more resilient to adversarial content than cheap ones?
- How do token-masking patterns distinguish genuine documents from poisoned ones?
- Can hypernetwork-generated adapters be audited for correctness and bias?
- Can deterministic programmatic checks prevent LLM hallucination in exploits?
- Can defenses tuned against appended attacks stop prepended payloads?
- Does the A-I-R framework distinguish insider attacks from adversarial positions?
- How does phase-awareness prevent false positive exploit paths in static analysis?
- How many probes does an attacker need to reach near-zero classification error?
- What feedback signal lets an attacker learn response distributions during classification?
- What makes semantic attacks harder to defend against than algorithmic ones?
- Do legitimate task signals exploit the same position and framing vulnerabilities as attacks?
- What distinguishes flow-preserving measurement from cognitive vulnerability profiling?
- Why do Claude and OpenAI models cheat through different strategies?
- Can false positives from input filtering be reduced without sacrificing defense?
- Did the attacker's framework use the same LLM model as defenders?
- Does naming a specific hack in prompts prevent only that hack or broader classes?
- What makes dense retrievers vulnerable to partition-based poisoning exploitation?
- How does context grafting perform on the same thirty-three failed runs?
- Why does scanning skill pairs not fully prevent cross-skill attacks?
- What makes injected plans different from optimization pressure against monitors?
- How does semantic framing differ from content injection attacks?
- What makes attractor-based probing better for third-party model auditing than alternatives?
- How do proxies stay faithful to real environments while reducing interaction cost?
- Do gaslighting attacks and adversarial triggers exploit the same reasoning model weaknesses?
- How do compress gates assume injection payloads appear at the user-prompt boundary?
- What defensive advantage does stigmergy offer over unmonitored channel analysis?
- Which of the 167 numerical features carry the strongest detection signal?
- What economic incentives make advertisement embedding attacks persistently viable?
- Can activation patching identify what a component encodes without representation?
- What feedback does ChainGuard return that an attacker could optimize against?