SYNTHESIS NOTE
Topics›Flaws›this note

Do frontier LLMs silently corrupt documents in long workflows?

DELEGATE-52 tests whether state-of-the-art language models reliably preserve document integrity across extended delegated tasks. Understanding this matters because single-step benchmarks may mask compounding failures that emerge only at workflow scale.

Synthesis note · 2026-05-18 · sourced from Flaws

Delegation requires trust — the expectation that an LLM will execute a task without introducing errors. DELEGATE-52 stress-tests that expectation with 310 work environments across 52 domains (coding, crystallography, music notation, genealogy) and a round-trip relay protocol where each task is paired with its inverse, so a perfect model would recover the original document exactly.

Across 19 LLMs, even frontier systems (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) corrupt an average of 25% of document content by the end of long workflows. Weaker models fail more severely. The degradation curve decelerates but does not plateau — the first half of an extended relay accounts for 2-3x more loss than the second half, yet the strongest model still drops below 60% accuracy by round-trip 50. Distractor files, longer documents, and longer interactions all worsen the rate.

The structural problem: errors are sparse but severe and they compound silently. A user reviewing one or two outputs sees competent work. A user delegating an end-to-end workflow gets a document that looks intact but contains accumulated drift in places they did not check. The trust assumption that holds at single-step interaction collapses at the timescale where delegation is actually valuable.

This is not a "weak model" finding. It is a ceiling on delegated work at the current frontier — one that scales unfavorably with exactly the workflow length that makes delegation attractive.

Inquiring lines that read this note 145

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What happens to knowledge when intelligence becomes tokenized like a commodity? Why do token-level mechanisms matter for learning to reason? Why do locally safe actions create system-level safety gaps? What causes retrieval-augmented generation systems to fail despite access to external knowledge? When do semantic similarity approaches miss structural retrieval failures? What compositional reasoning failures limit large language models despite scale? How does the generation-verification gap limit what we can measure about AI reasoning? Why don't LLMs reliably translate capability into accurate outputs? How do evaluation practices shape which failures stay visible? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? Can compression size predict model complexity better than parameter count alone? How do multi-agent LLM systems fail distinctly compared to single agents? Can parallel reasoning outperform sequential reasoning under fixed token budgets? What execution architectures enable agents to most effectively use tools? What mechanisms preserve shared understanding in evolving conversations? Do reasoning benchmarks predict model performance in long-horizon workflows? Do knowledge graphs offer advantages over embeddings for multi-hop retrieval? What do systematic disagreements between annotators reveal about ground truth? What training data selection strategies maximize generalization across difficulty levels? Do language models respond to social pressure and face-saving like humans? Do language models reason like humans or mimic surface patterns? What trajectory-level metrics beyond task success best evaluate agent performance? What prevents conversational agents from taking initiative in dialogue? How can we prevent synthetic data from contaminating statistical inference and corpora? Why can't prompting alone inject genuinely new knowledge into models? How effectively can language models perform reasoning, especially combined with symbolic methods? How do capability benchmark scores systematically misrepresent true model abilities? How can we distinguish genuine model deception from honest errors? What determines appropriate intervention timing and manner for AI agents? How do prompting refinements mask underlying biases and model frequency patterns? How can infrastructure records verify actual agent behavior? How do standardized protocols improve multi-agent coordination and reliability? Do reasoning traces faithfully reflect actual model reasoning? Can harness architecture and protocols provide agent reliability without model scaling? Do multi-agent systems introduce security vulnerabilities that single-agent architectures avoid? Why does memory consolidation cause performance regression in continual learning? Can memory architectures handle ultra-long context better than attention? How does misalignment propagate through agent communication networks? How should agents manage memory granularity to improve long-term performance? Can local safety checks guarantee system-level behavioral safety? How does self-revision in reasoning models affect accuracy and confidence? How does decomposing tasks improve reasoning and prevent failure propagation? Can self-generated feedback reliably guide model training without ground truth? What reasoning architectures enable models to solve complex problems efficiently? Does alignment training create genuine alignment or just output compliance? What fundamental constraints limit how effectively agents can improve themselves? Why does polished presentation create unearned authority in AI outputs? Why is hallucination an inevitable limitation of current language models? Why does adding new knowledge through fine-tuning degrade existing capabilities? How can we detect and prevent harm propagation through multi-agent delegation workflows? What capability trade-offs arise from domain specialization through fine-tuning? How do surface patterns enable correct outputs but reduce robustness? How does harness optimization generalize across different model architectures and domains? Can validator consensus certify semantic correctness beyond agreement? How do coordinated agents balance protocol compliance with reward maximization? How should agent systems validate and persist generated code artifacts?

Related concepts in this collection 6

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
15 direct connections · 138 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

frontier LLMs silently corrupt 25 percent of document content over long delegated workflows without plateauing