Explaining AI Agents Through Execution Traces
AI Agents are increasingly deployed in real-world settings, where they interact with external tools and make sequential decisions with limited human oversight. This creates a pressing need for reliable and auditable explanations of what an agent did and why. However, traditional Explainable AI (XAI) methods fall short of providing the process-level transparency required for such interactive, multi-step systems, motivating a paradigm shift toward approaches specifically designed for AI Agents. To address this gap, we present a post-hoc XAI framework that transforms a lengthy agent’s execution trace into a structured report and a faithful natural-language explanation explicitly grounded in its observable behavior. Because it relies solely on execution traces, the framework applies across different agent architectures, environments, and tasks. Human and automated evaluations across multiple benchmarks and architectures show that our framework produces highquality, trace-faithful explanations while reliably identifying unsupported claims, unjustified actions, and evidence gaps, outperforming naive LLM-generated explanations.
Introduction. Recent advances in Agentic AI are reshaping the capabilities of AI systems by enabling autonomous goal pursuit, adaptive decision-making, and complex task execution with minimal human oversight (Acharya, Kuppan, and Divya 2025). Unlike traditional AI models, AI Agents1 interact with external tools, maintain internal state, coordinate multiple actions, and adapt their behavior over extended execution horizons. While these capabilities unlock significant opportunities across critical domains, they also introduce new challenges for transparency, accountability, and governance (Zhu et al. 2026; Shah et al. 2026), making surface-level transparency no longer sufficient to support user trust or regula- tory compliance (Khalid, Farooqi, and Bilal 2026). Recent studies further highlight novel failure modes, including error cascades, responsibility gaps, flawed execution monitoring, and multi-step error propagation that obscures the origins of system behavior and outcomes (Zhu et al. 2026; Shah et al. 2026), highlighting the need for a rethinking of current XAI practices.
Lines of inquiry this paper opens 24
Research framings built by reading the notes related to this paper — the questions it feeds into.
Why do agents falsely report success on failed tasks?- Can explanations grounded in observable behavior recover an agent's internal reasons for acting?
- What tasks do AI agents still fail at most often?
- Can an auditor verify environment state without trusting the executor's self-report?
- Can execution traces reveal unsupported claims in AI agent behavior?
- Why do some occupations need human-AI partnership more than others?
- What task characteristics determine whether humans or agents should handle work?
- Which AI capabilities matter most for human-facing deployment contexts?
- How does capability differ from what workers actually want from AI?
- What happens to human bargaining power when interpersonal skills become the only remaining labor?
- What economic role remains for human labor after bottleneck automation?
- Does parallel task structure determine optimal multi-agent architecture?
- How does distributed coordination fail as agent networks scale?
- What capability threshold do agents need to self-organize effectively?
- Does horizontal coordination improve with stronger individual agents?
- At what capability threshold does multi-agent coordination stop helping?