Do frontier LLMs silently corrupt documents in long workflows?
DELEGATE-52 tests whether state-of-the-art language models reliably preserve document integrity across extended delegated tasks. Understanding this matters because single-step benchmarks may mask compounding failures that emerge only at workflow scale.
Delegation requires trust — the expectation that an LLM will execute a task without introducing errors. DELEGATE-52 stress-tests that expectation with 310 work environments across 52 domains (coding, crystallography, music notation, genealogy) and a round-trip relay protocol where each task is paired with its inverse, so a perfect model would recover the original document exactly.
Across 19 LLMs, even frontier systems (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) corrupt an average of 25% of document content by the end of long workflows. Weaker models fail more severely. The degradation curve decelerates but does not plateau — the first half of an extended relay accounts for 2-3x more loss than the second half, yet the strongest model still drops below 60% accuracy by round-trip 50. Distractor files, longer documents, and longer interactions all worsen the rate.
The structural problem: errors are sparse but severe and they compound silently. A user reviewing one or two outputs sees competent work. A user delegating an end-to-end workflow gets a document that looks intact but contains accumulated drift in places they did not check. The trust assumption that holds at single-step interaction collapses at the timescale where delegation is actually valuable.
This is not a "weak model" finding. It is a ceiling on delegated work at the current frontier — one that scales unfavorably with exactly the workflow length that makes delegation attractive.
Inquiring lines that read this note 145
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What happens to knowledge when intelligence becomes tokenized like a commodity? Why do token-level mechanisms matter for learning to reason?- How does token generation as flow differ from print's archival storage?
- Why does token ordering in LLMs create sequences rather than true temporal flow?
- How do current safety benchmarks miss pragmatic alignment failures?
- Why can every step pass its local check while a workflow still fails?
- Do Doc2Query approaches suffer from the same misaligned-target problem?
- Can semantic query expansion overcome vocabulary mismatch in corrupted text?
- What makes draft-centric systems better anchors for coherence than feed-forward outputs?
- What detection mechanisms work best for corruption-style document errors?
- Can detection mechanisms like diff review catch corruption better than deletion?
- Why do intermediate LLM layers become more precise in frontier models?
- What specific information must be exported from the language system?
- How do dependency errors propagate through incorrectly formalized definitions?
- Can evaluators investigate dependencies without accumulating mistakes over time?
- How does the rate of generation outpace archival of outputs?
- Why do language models produce plausible outputs over accurate failure reports?
- What workflow structure pairs LLM generation with human evaluation most effectively?
- What distinguishes entity errors from relation errors in LLM output?
- Which use cases can tolerate unverified LLM outputs without external verification?
- Do standard language benchmarks underestimate what LLMs can actually do?
- Can auditing LLM performance on complex inputs improve NLP pipeline reliability?
- Can critique-only calls in LLMs exploit a measurable gap between generation and evaluation?
- What happens when we treat LLM outputs as sampled rather than stored?
- Why do backward-looking benchmarks underestimate LLM scientific value?
- What causes silent document corruption in long LLM workflows?
- Why do LLMs choose incorrect edits despite understanding the task?
- What structural differences between human and LLM production create detectable signatures?
- Can we systematically enumerate LLM failure modes from first principles?
- Can LLMs reliably audit other language models for errors?
- Can tool use or self-conditioning fix long-horizon delegation drift in LLMs?
- Why do LLM outputs need verification even when they look polished?
- How long does retrievability support error detection across repeated LLM use?
- How do autonomous pipelines identify and fix silent bugs in data pipelines?
- How does error avalanching differ from entropy collapse as a failure mode?
- How do traditional quality assurance methods fail for mutable AI outputs?
- What failure modes does the negative-space checklist generation method actually catch?
- Why do frontier model failures in document editing go undetected by users?
- Does refining around bad results risk cascading errors in automated research?
- Can automating failure absorption hide problems that governance needs to surface?
- What makes code inspectable feedback more reliable than natural language verification?
- How do workflows normalize and hide errors before they become visible hazards?
- What does it mean for errors to remain visible, contestable, and recoverable?
- Why does each rewrite cycle degrade domain-specific details differently than compression?
- Why do LLMs strip applicability conditions during memory abstraction?
- Can task-agnostic compression of documents remain broadly useful for later queries?
- Why do LLMs degrade on long inputs before hitting context limits?
- Why do naive pruning and quantization destroy LLM performance so easily?
- What specific network sizes trigger coordination degradation in LLM systems?
- Can smaller LLMs perform tool use tasks through modular decomposition?
- Can you compose independent LLM experts without synchronization overhead?
- What architectural changes would accelerate the cleanup phase?
- Why does pre-computed workflow generation work better than runtime tool discovery for data security?
- How do insert-expansions and third position repair together cover full repair lifecycle?
- How do insert-expansions differ from third position repair in timing?
- Can structured output formats reduce instruction following degradation?
- How does task contamination differ from test set data leakage?
- How should benchmarks evaluate workflow architecture versus raw model performance?
- Can tool use or self-conditioning fix degradation in extended LLM workflows?
- Do synthetic verification chains from long-CoT models match the quality of human-annotated process labels?
- Do reasoning benchmarks predict real performance in long delegated workflows?
- Which model capabilities actually matter for sustained workflow delegation?
- Do short interaction benchmarks predict how LLMs perform in long workflows?
- Why do sparse per-step errors accumulate undetected across delegated tasks?
- What extraction errors most reliably propagate through knowledge graph traversal?
- Can small edits to source text compromise entire knowledge graph reliability?
- Can measuring semantic entropy help us detect unreliable generations?
- How do local soundness signals work across different problem domains?
- Can curriculum degradation of document quality accelerate policy learning?
- What training data contamination rates threaten model safety most practically?
- How does trajectory filtering handle noise when language models use code execution tools?
- How do trajectory quality and memory hygiene differ as evaluation metrics?
- Can marking AI provenance solve the grounding problem for generated text?
- Can provenance tracking prevent synthetic content from polluting the corpus?
- How does fluent output mask the mythic function of a system?
- How should organizations redesign workflows if LLMs cannot solve optimization directly?
- What prevents monolithic LLMs from coordinating decomposition with execution?
- Why do text-only benchmarks underestimate deployed model capability?
- Can review effort alone keep pace with frontier model degradation?
- Why do static benchmarks miss frontier capabilities that open-world tasks reveal?
- What makes provenance infrastructure more critical than artifact quality?
- How can anchored records fail authenticity while passing integrity checks?
- Why do frontier models corrupt more documents than weaker models during workflows?
- Why does increased model capability make detection harder in delegated workflows?
- How does workflow scale change the failure modes of frontier models?
- How does model tier affect whether errors delete or corrupt document content?
- What degradation patterns emerge as relay length increases in delegated tasks?
- What four domain properties make self-healing failure loops actually work?
- Can end-to-end models maintain debuggability without modular components?
- How should versioning and rollback govern the fast scaffold update loop?
- Why do frontier models corrupt documents while weaker models delete them?
- Can memory consolidation fragility be detected and reversed during execution?
- Why does LLM memory consolidation regress below no-memory baselines?
- What makes a learned consolidation rule lossy and where does contamination enter?
- Why should consolidation be scheduled offline rather than during forward passes?
- Why does workflow position amplify malicious signals downstream?
- Why does workflow position amplify malicious signals in multi-agent relay chains?
- Which workflow positions concentrate the most downstream dependencies and influence?
- What governance risks emerge when agents communicate in unreadable text?
- Which workflow positions concentrate the most downstream dependencies?
- Does encoding governance into runtime loops scale as deployment environments become more complex?
- Why does credit assignment through memory rewriting avoid expensive LLM parameter updates?
- What specific failure modes emerge when agents retrieve stale or contaminated memories?
- Can workflow memory compound reusable skills into measurable success improvements?
- What makes a deployment paradigm credible for maintaining scientific integrity?
- Can human inspection of auto-generated workflows catch harmful or incorrect API compositions?
- Can a system pass all local checks while the overall workflow still fails?
- How does error accumulation in workflows scale across multiple model calls?
- How do composite workflow and recurring pattern skills differ from atomic operation skills in scope?
- Does bounding textual edits prevent skill degradation better than free rewriting?
- How do frozen executors with editable text state compare to end-to-end fine-tuning?
- Can delegation prevent silent corruption in long delegated workflows?
- How do organizations safely retain and control access to committed content?
- Does the paper treat storage traces as addressed messages or unmarked traces?
- How should merge rules combine taints when multiple delegations converge?
- Do chain-level and flow-level checks face the same copyable-policy problem?
- Which shared channels cause the strongest correlated validator failures?
- Can a quorum of protocol-compliant validators certify semantically invalid transitions?
Related concepts in this collection 6
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Does model capability change how documents degrade?
This explores whether weaker and frontier LLMs fail in fundamentally different ways when handling long-form document tasks, and whether that difference affects how reliably we can detect failures in practice.
same paper, mechanism for why frontier failure is harder to detect
-
Can better tools fix LLM document editing errors?
Does giving LLMs agentic tool access—like diffing, re-reading, or structured editors—improve their reliability on long-horizon document workflows? Understanding whether the problem is tool limitations or decision-making quality matters for reliability engineering.
same paper, fixes that do not work
-
Do short benchmarks predict how models perform over long workflows?
Standard LLM benchmarks measure single-turn performance, but real workflows involve sustained delegation across many turns. The question explores whether top benchmark performers maintain accuracy through longer interaction chains.
same paper, methodology implication
-
Do models fail worse when their own errors fill the context?
As a model's prior mistakes accumulate in context, does subsequent accuracy degrade predictably? And can scaling or architectural changes prevent this self-contamination effect?
adjacent mechanism for compounding error
-
Why do language models fail to act on their own reasoning?
LLMs produce correct explanations far more often than they produce correct actions. What causes this knowing-doing gap, and can training methods close it?
adjacent: capable rationale but unreliable execution
-
Can safety tests miss hazards that build over time?
Static tests check individual responses, but systems can accumulate unsafe state across interactions. This explores whether snapshot evaluations are sufficient to catch hazards that emerge only through repeated use or stored context.
grounds: independent evidence for the temporal shape, sparse per-step errors that accumulate over a workflow so short tests look clean; consistent with that claim, does not test it
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- LLMs Corrupt Your Documents When You Delegate
- The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs
- LLMs Get Lost In Multi-Turn Conversation
- Emergent Collusion in Long-Horizon LLM Agent Interaction
- FrontierChallenge: Evaluating Scientific Workflow Completion
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
- Can You Trust LLM Judgments? Reliability of LLM-as-a-Judge
- Trust propagation and structural containment in Multi-agent LLM pipelines
Original note title
frontier LLMs silently corrupt 25 percent of document content over long delegated workflows without plateauing