Stealing Reasoning Traces from Proprietary LLM APIs

Paper · arXiv 2608.09867 · Published August 10, 2026
LLM Failure Modes

Leading large language model providers now conceal their models’ step-by-step reasoning, or chainof-thought, to protect intellectual property and limit information leakage. Rather than storing these traces server-side, providers return them to the client as blocks of encrypted text, which the client passes back with each subsequent request. Building on prior research, we identify an architectural vulnerability: these encrypted blocks are fully compatible and interchangeable across different sessions, users, and models within a provider’s ecosystem. We exploit this compatibility to develop a scalable decryption jailbreak. By injecting an encrypted reasoning trace from a given model into a weaker, and less safeguarded model from the same provider, we force it to decode and output the trace verbatim in plaintext, without ever jailbreaking the more capable model directly. This vulnerability enables four distinct attack vectors. First, it circumvents anti-distillation mechanisms, allowing adversaries to extract a proprietary model’s reasoning, as we demonstrate across Anthropic, OpenAI, and Google. Second, it allows for large-scale private data extraction. Developers frequently share session logs publicly, unaware of contents of the encrypted blocks.

Introduction. Frontier large language models have increasingly evolved into “reasoning models”. Before producing a response visible to the user, these models generate extensive internal chains of thought – a technique that has driven substantial leaps in performance and complex problem-solving (Jaech et al., 2024). However, these hidden traces act as an internal monologue that often contains far more dense and sensitive information than the final output, including intermediate hypotheses, tool outputs, user data, and contextual secrets. Exposing these reasoning traces in plaintext leaves proprietary systems highly vulnerable to model distillation by competitors (Muennighoff et al., 2025), and it risks unmasking internal safety and refusal mechanisms or revealing harmful information (Green et al., 2025; Mao et al., 2026).

Discussion / Conclusion. The transition toward reasoning models has introduced new complexities in balancing intellectual property protection with system security. While current API designs utilize client-side encrypted reasoning blocks to mitigate server storage costs, our research demonstrates that the broad cross-compatibility of these blocks creates unintended decryption channels, enabling model distillation and other attacks. Looking forward, as these models are increasingly integrated into complex workflows, they will inevitably process growing volumes of private and sensitive user data. This intersection of pervasive data collection and encrypted, illegible reasoning introduces critical challenges for the future of AI transparency. When models utilize sensitive data – such as personal information or API keys – to make decisions within a hidden chain of thought, users lose visibility into how their information is being processed.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Can reasoning traces and behavior monitoring reliably detect hidden AI scheming? What attack surfaces do reasoning traces and chains introduce? What drives appropriate trust calibration in personalized AI systems? Why do people disclose to AI systems despite their artificial nature? How does persona conditioning amplify demographic stereotyping and bias in models? Do reasoning traces faithfully reflect actual model reasoning? What causes retrieval-augmented generation systems to fail despite access to external knowledge? What safeguards enable trustworthy AI-assisted scientific peer review at scale? Can reasoning scale in latent space without tokens? Why can recurrent transformers achieve reasoning capabilities that standard transformers cannot? Why do some clarifying approaches produce understanding while others just satisfy? How well do AI systems understand human social norms? Do language models learn genuine understanding or just surface patterns?