Featured

Latent Reasoning with Normalizing Flows

Guancheng Tu, Xiangjun Fu, Suhao Yu, et al. · arXiv:2606.06447

The tension between explicit and implicit reasoning in language models has sharpened considerably: while machines can reason through hidden latent computation instead of generating visible chains, most deployed systems still rely on verbalized chain-of-thought because it integrates naturally with autoregressive generation and probabilistic inference. This paper enters that conversation by asking whether normalizing flows—which can model compact continuous distributions tractably—might unlock latent reasoning without sacrificing the practical advantages (KV caching, left-to-right sampling, likelihood estimation) that make explicit CoT appealing at scale. The core insight is architectural: rather than choosing between semantic reasoning in continuous space and computational efficiency in discrete token streams, NF-CoT attempts to have both by interleaving normalizing-flow generations (for thoughts) with standard LM-head generations (for text) in the same causal pipeline. Yet the framing leaves open a deeper question about what makes reasoning "latent" useful in practice—whether the efficiency gains come primarily from compressing intermediate steps, or whether there's something about reasoning in continuous space that fundamentally changes how models explore the solution landscape.

Abstract

Large language models often improve reasoning by generating explicit chain-of-thought (CoT), demonstrating the importance of intermediate computation. However, textual CoT forces this computation through a discrete, serial, and communication-oriented token stream: each reasoning step must be verbalized before the model can proceed, even when the underlying update is semantic, uncertain, or only partially formed. Latent reasoning offers a higher-bandwidth alternative by performing intermediate computation in compact continuous states before committing to text. Yet existing latent-reasoning methods often sacrifice key advantages that make CoT effective in autoregressive language models, including native left-to-right generation, probabilistic sampling, compatibility with KV-cache decoding, and tractable likelihood estimation. We propose NF-CoT, a latent reasoning framework that preserves these advantages by modeling continuous thoughts with normalizing flows. NF-CoT instantiates a TARFlow-style normalizing flow inside the LLM backbone, defining a tractable probability model over compact continuous thoughts distilled from explicit CoT. Continuous-thought positions are generated by an NF head, while text positions are generated by the standard LM head within the same causal stream. This design provides exact likelihoods for latent thoughts, enables probabilistic left-to-right decoding with the original KV cache, and supports direct policy-gradient optimization in the latent reasoning space. On code-generation benchmarks, NF-CoT improves pass rates over explicit-CoT and prior latent-reasoning baselines while substantially reducing intermediate-reasoning cost.

Synthesis notes nearest this paper, framed as questions — click to read.

Explore in faceted view

Not questions with answers — ways of approaching this research. Each opens a synthesized line of inquiry across the collection.


All featured →