Does reasoning ability actually degrade with longer inputs?
Explores whether modern language models can maintain reasoning performance when processing long contexts, and whether technical capacity translates to practical reasoning capability over extended text.
The FLenQA benchmark exposes a critical gap between technical context window capacity and actual reasoning capacity over long inputs. By embedding simple reasoning tasks (True/False questions requiring integration of two information pieces) within irrelevant padding text of varying lengths, the paper shows that reasoning accuracy drops from 0.92 to 0.68 at just 3000 tokens — far below any modern model's context window.
Three findings make this particularly concerning:
1. The degradation is task-agnostic. Regardless of whether padding text is similar or dissimilar to the reasoning content, and regardless of where the information pieces are embedded within the context, similar degradation trends appear. The failure is not about content interference but about attention dilution over length.
2. Next-word prediction performance is uncorrelated with reasoning performance. Models that maintain strong perplexity on long inputs still fail at reasoning over those inputs. This means language modeling benchmarks on long contexts are misleading indicators of actual long-context utility — a model can "understand" the text (predict tokens well) while failing to reason over it.
3. CoT does not mitigate proportionally. Chain-of-thought prompting increases accuracy roughly uniformly across context lengths but does not close the length-induced gap. The degradation persists under CoT because the bottleneck is in information retrieval from context, not in reasoning over retrieved information.
This is a complementary mechanism to Why do language models fail at temporal reasoning in complex tasks?. That failure is about task complexity; this is about input noise. Together they define a two-dimensional reliability surface: reasoning degrades with both task complexity AND input length, and the two dimensions are independent.
The implication for RAG systems is direct: retrieved documents add to input length, and if that length includes irrelevant passages (as it typically does), reasoning over the retrieved content degrades even when the relevant information is present. Since Why does vanilla RAG produce shallow and redundant results?, the length degradation explains part of why static retrieval fails — more retrieved documents means more padding means worse reasoning.
A complementary training-time finding complicates this picture. "Longer Context, Deeper Thinking" (2025) shows that models with stronger long-context capacity (128k vs 32k) consistently achieve higher accuracy on mathematical reasoning benchmarks (MATH500 and AIME) — even when test-time inputs are short. Long-context training benefits reasoning as a foundation, not just for processing long inputs. The implication: the inference-time degradation documented in this note coexists with a training-time benefit. Models trained on longer contexts develop better reasoning foundations, but at inference time, longer inputs still degrade performance. The two findings are compatible: long-context training may improve the base reasoning capability, while inference-time input length introduces the noise and distraction effects that degrade it. Source: Arxiv/Evaluations.
Inquiring lines that read this note 171
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Does AI assistance erode cognitive skills while inflating perceived competence? How do interpretive frames override surface features in text comprehension?- How do readers selectively hold frame-related words in mind?
- Can adding more words to a passage actually interfere with meaning?
- Does high knowledge density in text reduce user motivation to read more?
- Can input augmentation and rephrasing compensate for smaller model limitations?
- Does irrelevant content degrade reasoning even when it fits the context window?
- Can manipulative prompts reduce reasoning model accuracy without fine-tuning?
- Does irrelevant context degrade reasoning even within model context limits?
- How should reasoning prompts adapt based on question complexity and type?
- What prompting strategies most effectively boost long-context LLM performance on retrieval?
- How do smaller models respond to longer reflection prompts?
- Can structured prompts reduce reasoning steps while improving financial accuracy?
- Can operationalizing theory into prompt structure improve reasoning more than theory itself?
- How do input length and context size separately affect reasoning quality?
- Does sentence-level granularity capture enough structure for complex reasoning tasks?
- How does SONAR embedding quality affect downstream reasoning accuracy?
- Where do humans and language models actually diverge in reasoning ability?
- What causes snowball errors to accumulate across reasoning steps in language models?
- Can long-context readers handle compositional tasks or just semantic search?
- What neuroscience evidence suggests language networks are not optimized for reasoning?
- How can entailment benchmarks separate genuine reasoning from memorization effects?
- Why does extended reasoning fail for search and knowledge retrieval tasks?
- Can long-context models handle compositional reasoning requiring structured logic?
- What makes deductive reasoning so brittle in language models overall?
- How much does schema bloat actually degrade reasoning in large language models?
- Can simple structure perturbations reliably expose memorization in reasoning models?
- Do distributed relational tasks consistently underperform local classification across NLP domains?
- Can cognitive scaffolding replace tool-based reasoning augmentation in language models?
- Why do long-context language models struggle with compositional reasoning tasks?
- When is numeric computation the real bottleneck versus reasoning depth?
- How does recombining partial trajectories maintain coherence in natural language reasoning?
- How does evidence retrieval affect compositional reasoning in language models?
- How do depth and length constraints emerge in latent reasoning without explicit training?
- What distinguishes genuine reasoning activation from memorization-assisted answer recall?
- Can latent reasoning architectures work as retrofits to existing models?
- How much does pre-training frequency predict reasoning task performance?
- Can models trained on longer contexts develop better fundamental reasoning abilities?
- Do base models truly possess latent reasoning capability?
- Why do reasoning tasks improve more than retrieval from lookup memory?
- Can auxiliary modules preserve reasoning without catastrophic forgetting?
- What kinds of reasoning tasks reveal the ceiling of text-only training?
- Can meaning-level metrics like Semantic Entropy avoid length bias?
- Why do large language models fail at temporal reasoning in complex legal cases?
- Does more thinking always help large language models or sometimes hurt?
- How does the distance between natural language and formal notation affect translation accuracy?
- Can context compression preserve what matters without introducing bias?
- How does separating local and global context dependencies affect long-context performance?
- Does recurrent memory or gist compression work better for ultra-long context?
- Can recurrent state mechanisms process longer sequences than attention-based working memory approaches?
- How does externalized state affect the long-context bottleneck in language models?
- How does reducing activation precision further extend context length?
- How do recurrent memory systems handle ultra-long context differently than attention?
- What capacity limits does the memory model face as corpus grows?
- How does context length affect retrieval quality in modernized BERT architectures?
- What happens to anaphoric reference when context exceeds the window?
- Why do longer context windows alone fail to capture temporal dynamics in dialogue?
- What makes a background condition relevant to a specific reasoning task?
- Why do models automatically adjust reasoning length to problem difficulty?
- Do reasoning models trade instruction following for deliberative capability?
- Does model scaling improve knowledge storage faster than reasoning ability?
- Does reasoning structure match explicit versus implicit task demands?
- Why do reasoning models fail when input length increases even below context limits?
- How does scaling reasoning capability actually reduce instruction-following ability?
- What changes when reasoning models adopt trajectory-response output formats?
- What causes reasoning quality to degrade during long research tasks?
- Why might rationales that predict common text patterns fail on hard novel reasoning?
- Is reasoning failure caused by task complexity or training distribution gaps?
- Why does reasoning performance degrade as input length increases?
- Can weak models reason better when freed from cognitive load by structure?
- Does longer reasoning always improve model accuracy on complex tasks?
- Why does long-form generation need different retrieval than factoid questions?
- Why do longer queries benefit less from clarification questions?
- Why does explicit reasoning degrade passage reranking performance?
- When does long-context LLM reasoning fail where structured retrieval succeeds?
- Does filtering passages before generation improve large model answer quality?
- How do retrieval heads interact with layer-level separation of knowledge and reasoning?
- What is the optimal balance between search rounds and reasoning depth per round?
- What makes active reasoning through dialogue harder than passive reasoning?
- How does structural complexity in sentences degrade LLM reasoning systematically?
- How do logic units preserve document structure better than fixed-size chunking?
- Why do format and structure matter more than actual content in reasoning?
- Why does scheme classification require more cognitive load than identifying premises?
- Can marginal hints integrate better into reasoning than comprehensive explanations?
- Does explicit reasoning help or hurt tasks requiring continuous nuanced judgment?
- How should iterative research tasks limit context per reasoning turn?
- Can extended reasoning training capture individual strategic thinking styles?
- How does random walk length control reasoning complexity in question generation?
- Why might latent reasoning capture types of thinking that verbalized CoT cannot?
- Why do longer reasoning chains signal hesitation rather than depth?
- Can latent reasoning mechanisms and recursive tracking mechanisms be combined effectively?
- Can dataset design systematically expand reasoning graph diameter?
- How much reasoning depth do we actually need for most real-world tasks?
- Can minimal reasoning steps match verbose reasoning accuracy?
- Can bounded workspaces prevent overthinking better than summarization alone?
- What computational structures can actually scale serial reasoning depth?
- How does era sensitivity in legal cases compound with context length failures?
- When should an LLM engage extended reasoning versus responding directly?
- How does context complexity affect LLM performance on temporal reasoning tasks?
- Why does premise ordering shift syllogistic reasoning performance by over 30 percent?
- Why do correct reasoning traces in language models tend to be shorter?
- Can concise reasoning traces match verbose explanation accuracy?
- Why do temporal reasoning patterns matter more than final answers?
- Can derivational traces be distinguished from stylistic mimicry of reasoning?
- Do shorter reasoning chains maintain instruction adherence better than longer ones?
- Why does concise reasoning maintain accuracy with far fewer tokens?
- What quality filters distinguish useful reasoning enrichment from shallow repetition?
- Do tokens beyond a critical threshold actually improve reasoning quality?
- Does thinking-token overuse actually degrade reasoning accuracy in practice?
- Does more thinking always improve language model accuracy?
- Are larger models and search access substitutes for factual accuracy?
- How do frontier models maintain agreement scores above 90 percent across reasoning tasks?
- Should long-context evaluation measure the coupled system?
- Why does training data format shape reasoning strategy more than domain content?
- Does training data format shape reasoning strategy more than domain content?
- How much does training data format influence reasoning strategy versus domain content?
- Why do language models fail at pronouns across distant segments?
- Why do language models fail at coreference across long contexts?
- Can benchmark performance distinguish surface from structural linguistic knowledge?
- Do pretrained language models carry reusable computational scaffolding for length handling?
- Why do thinking models execute longer tasks than standard language models?
- Can autoformalisation from natural language preserve semantic accuracy?
- How does tool-based reasoning expand what language models can do?
- Why do some models like Llama degrade under long context?
- Can episodic and semantic memory improve long-horizon task reasoning?
- How do adaptive memory modules compare to feedback-based working memory for long context?
- How does implicit meaning processing limit LLM pragmatic reasoning?
- Can language models reason without relying on surface level pattern matching?
- Can language models perform genuine symbolic reasoning without semantic grounding?
- Does more inference compute help reasoning models match specialized domain performance?
- Can post-thinking compute on memory reduce query-time reasoning costs?
- How do neural memory modules extend context length beyond attention limits?
- Why does attention quality degrade as context length increases?
- Could real-time search systems avoid era sensitivity in legal reasoning?
- Why do fixed-size document chunks break complex procedural question answering?
- Can long-context models replace retrieval-augmented generation systems?
- What structural properties define effective long chain-of-thought reasoning?
- Why does chain-of-thought prompting fail to fix length-induced reasoning degradation?
- How do longer reasoning chains create vulnerability to attacks?
- How does chain-of-thought length affect attention to constraint tokens?
- Why do longer reasoning chains correlate with lower accuracy in o1-like models?
- Why do concise reasoning chains match verbose chain-of-thought token efficiency?
- Does chain-of-thought accuracy degrade with longer reasoning traces?
- What makes specific-facet questions outperform generic need-rephrasing requests?
- Why do current speech benchmarks fail to measure reasoning over audio?
- Why does document perplexity stay low while question-answering accuracy drops?
- How do behavioral differentiation and paraphrase stability trade against accuracy?
- How do we measure genuine reasoning inside a language model?
- Why does representation recycling of MI-peak tokens improve reasoning accuracy?
- Can standard next-token prediction capture complex multi-step human reasoning directly?
- How do prior errors in reasoning context amplify future mistakes?
- How do prior errors in context history amplify future mistakes in long tasks?
- Does sequence length affect sparsity tolerance the same way across task types?
- What makes sparse attention more reliable for long-context retrieval?
- How does parametric knowledge sabotage context-grounded question answering?
- How do language models treat injected evidence as shared background knowledge?
Related concepts in this collection 5
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Why do language models fail at temporal reasoning in complex tasks?
Language models correctly answer simple temporal questions but produce logically impossible timelines in complex legal documents. This explores what task features trigger reasoning failures and whether the competence is genuinely lost or masked by surface-level patterns.
complementary failure axis: task complexity vs input length
-
Why does vanilla RAG produce shallow and redundant results?
Standard RAG systems get stuck in a single semantic neighborhood because their initial query determines what documents are discoverable. The question asks whether fixed retrieval strategies fundamentally limit knowledge depth compared to iterative exploration.
RAG retrieval adds length; length degrades reasoning
-
Does more thinking time actually improve LLM reasoning?
The intuition that extended thinking helps LLMs reason better seems obvious, but what does the empirical data actually show when we test it directly?
another dimension where "more" (tokens) ≠ "better" (reasoning)
-
Can long-context models resolve retriever-reader imbalance?
Traditional RAG systems force retrievers to find precise passages because readers had small context windows. Do modern long-context LLMs change what architecture makes sense?
challenges the long-context solution: reader burden increases with length but reasoning degrades
-
Do vector embeddings actually measure task relevance?
Vector embeddings rank semantic similarity, but RAG systems need topical relevance. When these diverge—as with king/queen versus king/ruler—does similarity-based retrieval fail in production?
compounds the length problem: semantic retrieval returns associated-but-irrelevant documents, creating exactly the irrelevant padding that FLenQA shows degrades reasoning; imprecise retrieval directly produces the input-length degradation
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Same Task, More Tokens: the Impact of Input Length on the Reasoning Performance of Large Language Models
- Longer Context, Deeper Thinking: Uncovering the Role of Long-Context Ability in Reasoning
- Think Deep, Not Just Long: Measuring LLM Reasoning Effort via Deep-Thinking Tokens
- Self-Guided Test-Time Training for Long-Context LLMs
- When More is Less: Understanding Chain-of-Thought Length in LLMs
- On the Reasoning Capacity of AI Models and How to Quantify It
- Between Underthinking and Overthinking: An Empirical Study of Reasoning Length and correctness in LLMs
- Beyond Accuracy: Evaluating the Reasoning Behavior of Large Language Models -- A Survey
Original note title
reasoning performance degrades with input length even far below context window limits