Line of inquiry
Inquiring lines›How do we keep AI systems safe and…›How do architectural choices affec…›this line of inquiry
Why do standard evaluation practices obscure safety-critical AI failures?
A broader line of inquiry — a family of 48 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 48
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Which evaluation habits keep safety-critical failures hidden in AI systems?
- What does it mean for errors to remain visible, contestable, and recoverable?
- Why do evaluation habits hide safety-critical challenges from view?
- How do response-centered evaluation assumptions hide safety-critical failure modes?
- What conditions allow technical systems to escape critical evaluation?
- What makes a model's errors visible and contestable to users?
- Can AI outputs inspire new directions even when they seem like failures?
- Why can't individual companies detect all emerging patterns in AI failures?
- How do inherited evaluation habits obscure failures that matter most?
- What would it take to measure whether system errors stay visible and contestable?
- Can users distinguish between automation errors and unauthorized agent actions?
- Can a correct outcome hide a fundamentally unsound decision-making process?
- Can external process logs make AI errors verifiable and harder to hide?
- How can a single instrument measure errors across multiple system layers?
- Can a system pass all local checks while the overall workflow still fails?
- Why do firms delay disclosing reliance on AI after errors surface?
- Can monitors fail together through shared training data or infrastructure?
- Does component-level checking detect system-level failures in pipelines?
- How do workflows normalize and hide errors before they become visible hazards?
- How does laboratory generalization evidence connect to deployment failure modes?
- Why don't users push back when AI makes obvious mistakes about false claims?
- What makes the frame problem distinct from feature-level shortcuts?
- Why haven't labs adopted self-hacking approaches to catch specification errors?
- Why does greater automation actually obscure rather than eliminate research failure modes?
- Why do quiet failures reach deployment scale more often than loud ones?
- What specific failure modes must evaluation catch before deploying action-capable systems?
- Can orchestration platforms and better infrastructure reduce AI correction time?
- What types of social situations cause all AI models to fail in identical ways?
- How does automation obscure failure modes in ways that make detection harder?
- Can automating failure absorption hide problems that governance needs to surface?
- Why is error rate alone misleading without strong contestability conditions?
- How do autonomous pipelines identify and fix silent bugs in data pipelines?
- What happens when students encounter errors they cannot resolve through prompting alone?
- How do past research mistakes prevent future pivot loops from repeating them?
- What design principles prevent error cascades in multi-step evaluation systems?
- Why did the same CoT exposure mistake occur at both OpenAI and Anthropic?
- Can human inspection of auto-generated workflows catch harmful or incorrect API compositions?
- What happens when monitors themselves become targets for optimization?
- What does recovery look like as a formal part of AI design?
- What makes some AI failures feel like regret instead of mistakes?
- Why did the UN panel treat AI failures as alignment problems instead of corporate misbehavior?
- How does implicit influence differ from omission in monitoring failures?
- Why are closed AI systems harder to hold accountable than open ones?
- What distinguishes an error bound from a forecast of system behavior?
- What distinguishes a component failure from a monitoring coverage failure?
- Why do one-shot transparency studies miss the temporal reversal entirely?
- What role does opacity play in preventing effective challenge to system outputs?
- How does generative variability intensify the problem of passive AI systems?