Line of inquiry
Inquiring lines›How do language models learn and r…›How do language models learn and w…›this line of inquiry
What capabilities differentiate diffusion from autoregressive language models?
A broader line of inquiry — a family of 49 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 49
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- Do diffusion language models learn differently than autoregressive models?
- How can diffusion models predict future tokens without completing prior blocks?
- What structural differences between diffusion and autoregressive models enable bidirectional prompting?
- Can gradient-based control reach properties that autoregressive methods cannot?
- Why do autoregressive models fail at controlling syntactic structure and semantic content?
- Can autoregressive models learn faithful translation to logical representations without semantic loss?
- Can diffusion models condition on right context natively without special training for infilling?
- Can fast-slow separation improve both memory and generation in language models?
- How does selective looping in diffusion models differ from recurrence in autoregressive architectures?
- Can architecture changes and early stopping combine to close the diffusion inference gap?
- Can diffusion language models match autoregressive inference speed in practice?
- Why is reinforcement learning harder to apply to diffusion language models?
- How does the articulatory substrate explain direct speech-to-speech superiority over transcription pipelines?
- Does diffusion's control advantage come from speed gains or from architectural differences?
- What causes autoregressive generation to fail on out-of-corpus item identifiers?
- Why does bidirectional attention in diffusion models prevent KV cache reuse?
- Can diffusion models perform infilling and reverse generation as naturally as forward generation?
- How does tokenization toward corpus mean affect downstream output diversity?
- Can autoregressive models be trained to produce more cataphoric text?
- How do diffusion language models outpace autoregressive generation in speed?
- Do speech models learn the articulatory processes that produce acoustic signals?
- Can speech embeddings carry articulatory structure that text cannot?
- How does the discrete token bottleneck prevent gradient flow in language model control?
- Do speech encoders actually learn the physics of how vocal tracts produce sound?
- Does bidirectional attention improve language models as universal encoders?
- Can feature disentanglement in gesture synthesis generalize to completely unseen voice distributions?
- Can one streaming model handle turn-taking better than cascaded ASR-LLM-TTS?
- Can looped diffusion models outperform standard depth scaling at fixed parameters?
- Can decoder-only models become effective text encoders with training?
- Do decoder-only models have inherent architectural limits for non-sequential information?
- How does causal multimodal modeling differ from encoder-decoder architectures?
- Why do hybrid paradigms outperform pure autoregressive or pure diffusion approaches?
- Do bidirectional and any-order generation expose different parts of the joint distribution?
- What specific optimizations from LLM training transfer back to encoder models?
- Why do encoder models process document corpora more efficiently than decoder models?
- Do discrete tokenized modalities preserve information better than continuous embeddings?
- Why does articulatory probing predict SSL model performance better than phonetic probing?
- How do speech encoders learn articulatory physics without phonetic labels?
- What information does transcription destroy that direct speech-to-speech models preserve?
- Can skipping transcription reduce speech dialogue latency below 300 milliseconds?
- Can textual gradients generalize natural language feedback across computation graphs?
- Does direct speech-to-speech generation really eliminate transcription latency?
- Can articulatory inversion serve as a window into what speech models have learned?
- How much latency improvement comes from collapsing the speech pipeline?
- How does removing transcription change speech-to-speech generation latency?
- Why does transcription destroy prosodic information in speech processing?
- How do different speech encoder layers capture different types of gesture information?
- What information does transcription destroy that direct speech pathways preserve?
- What paired speech data is needed to train end-to-end models?