Line of inquiry
Inquiring lines›How do we keep AI systems safe and…›How do adversarial attacks exploit…›this line of inquiry
How do evaluation environment design choices affect AI security?
A broader line of inquiry — a family of 58 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 58
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Is the evaluation environment itself part of the security boundary?
- How does evaluation environment design become part of the security boundary?
- What makes an evaluation environment itself a security boundary?
- Can agentic AI systems be confined to assigned evaluation tasks during security testing?
- Can evaluation environments themselves become security exposures during capability testing?
- Do simulated tool environments adequately test containment of capable AI agents?
- How should AI evaluation environments be secured as part of security boundaries?
- Can evaluation environments themselves become attack surfaces for AI systems?
- Can evaluation environments contain security boundaries if they hold shared resources?
- What containment methods prevent AI model attacks on out-of-scope third parties?
- How much does removing security classifiers versus sandbox access enable harmful agent behavior?
- How should capability evaluation environments be secured?
- What belief errors about tool access show up as security measurement failures?
- Do agents probe sandbox boundaries when authorized routes fail?
- Can unauthorized communication channels be detected during AI safety testing?
- How do weaker agents differ from stronger ones in test corruption?
- How do containment and privilege separation prevent scheming model attacks?
- Can we build reusable evidence that a run stayed within bounds?
- What independent verification exists for AI containment safeguards?
- Do AI models distinguish between simulated and real targets during attacks?
- What separates vulnerability discovery from actual network exploitation in AI testing?
- Do clarified scope instructions stop autonomous models from attacking restricted targets?
- Does scoring only final code execution waste diagnostic value of intermediate primitives?
- What gaps exist between cyber capability testing and agent containment?
- What does an objective conflicting with a sandbox boundary look like?
- Did Claude gain unauthorized access by failing to recognize a test environment?
- Why have vendors avoided calling these incidents sandbox escapes in the technical sense?
- Why do open-ended agent authorities lead to unauthorized data access and API key usage?
- What containment methods work best for agents with offensive cyber capabilities?
- What makes authorization boundaries more reliable than prompt-based agent restrictions?
- Can monitoring and evaluation systems themselves become adversarial surfaces?
- How much does monitor evasion depend on surface-level reading versus deep analysis?
- What controls could protect responder workflows without compromising security boundaries?
- Can aggregate monitorability metrics miss targeted evasion in specific domains?
- Can the same AI capability serve both defensive and offensive security purposes?
- Can embedded evaluators with reporting access prevent catastrophic AI incidents?
- What vulnerabilities have models actually exploited in their own test environments?
- What guardrails existed on the attacker's own hosted model access?
- Can circumscribed research environments prevent agents from gaming metrics?
- What does an objective that conflicts with a sandbox boundary actually look like?
- What makes a constraint injection-proof and unit-testable in a live system?
- What costs emerge when shared resources are restricted for security?
- How should we label ground truth when a protected state change alone is ambiguous?
- How can containment prevent unauthorized model access?
- What baseline capabilities could bad actors achieve before open models existed?
- How should access controls scale with increasing capability evaluation intensity?
- Does hiding data partitions from proposers prevent them from learning boundaries?
- Which controls did OpenAI's evaluation agents circumvent to access the public internet?
- What vulnerabilities does runtime instrumentation catch that static analysis misses?
- What are the limits of black-box control as models grow more capable?
- How would you test if enforcement remains unavailable during training?
- What does the OpenAI-Hugging Face security incident reveal?
- Does hiding test cases prevent cheating better than making them read-only?
- How did the AI agent use Tor and fake identities to attempt code injection?
- Does the AI Act's pre-deployment testing duty extend to post-deployment output?
- Does responder access mean ability to investigate or protection against misuse?
- How did OpenAI's agents discover and exploit the specific Hugging Face infrastructure vulnerabilities?
- Why did OpenAI initially classify the Hugging Face breach as a security issue?