Theme of inquiry
How do adversarial attacks exploit vulnerabilities in AI safety monitoring?
A question within its area, explored through 6 lines of inquiry below — each a family of specific questions the research asks.
80 specific questions
- Can deliberately corrupted reasoning traces fool safety evaluation systems?
- How can simple prompt injection attacks extract reasoning trace content?
- Does optimizing against CoT monitors inevitably produce obfuscated reasoning?
- Why does obfuscating reward hacking reduce the reliability of trace-based monitors?
- Why does detector performance flip sign between different model architectures?
- Can activation-level monitoring catch hacks that leave no clean trace?
- Can harmful reasoning be planted through context without fine-tuning the model?
35 specific questions
- Can current cybersecurity benchmarks measure model exploitation risk?
- Why do vulnerability reproduction benchmarks miss real exploitation ability?
- How do frontier AI models currently score on measured cyber offense capability?
- Can benchmark scores alone prove a model's exploitation capability?
- How does evaluation of exploit capability differ from other dual-use AI measurements?
- How did unsolved ExploitGym tasks correlate with escalating risk-taking?
- How do frontier models exploit vulnerabilities in their own evaluations?
27 specific questions
- Does a planted honeypot catch all the hacks that actually matter?
- How do planted honeypots distinguish earned passes from coincidental task completion?
- Does a planted honeypot count the hacks that actually matter?
- Why do honeypot tasks reveal reward hacking better than standard benchmarks?
- Can planted hacks within tasks meet the reusability requirement?
- Do honeypot shortcuts in synthetic tasks measure the hacks that actually matter?
- Can automatic honeypot detection replace human judgment of agent shortcutting?
83 specific questions
- Does chain-level defense reduce but not eliminate attack success rates?
- How do hardened prompts defend against adversarial attacks in multi-agent systems?
- Does prompt hardening equally protect single and multi-agent web systems?
- When do multi-agent architectures create more attack surface than single-agent systems?
- How do multi-step exploitation chains make agent containment harder to achieve?
- How does payload exposure compare between single and multi-agent architectures?
- How does task division in multi-agent design affect security outcomes?
58 specific questions
- Is the evaluation environment itself part of the security boundary?
- How does evaluation environment design become part of the security boundary?
- What makes an evaluation environment itself a security boundary?
- Can agentic AI systems be confined to assigned evaluation tasks during security testing?
- Can evaluation environments themselves become security exposures during capability testing?
- Do simulated tool environments adequately test containment of capable AI agents?
- How should AI evaluation environments be secured as part of security boundaries?
29 specific questions
- Why do completion-oriented models systematically sacrifice privacy compliance?
- Can increasing reasoning steps make models leak more private information?
- How do agent privacy compliance and task success differ in evaluation?
- How do minimal-disclosure privacy contracts enable multi-dimensional agent evaluation?
- How does completion-oriented bias in agents lead to unintended personal data disclosure?
- Can minimal privacy boundaries generalize beyond phone-use contexts?
- Can differential privacy during generation eliminate leakage at scale?