SYNTHESIS NOTE
Topics›Flaws›this note

Does model capability change how documents degrade?

This explores whether weaker and frontier LLMs fail in fundamentally different ways when handling long-form document tasks, and whether that difference affects how reliably we can detect failures in practice.

Synthesis note · 2026-05-18 · sourced from Flaws

DELEGATE-52 surfaces an under-discussed asymmetry in how LLM document degradation looks at different capability tiers. Weaker models fail loudly: they delete content. The document gets visibly shorter, sections disappear, structure breaks. A reviewer notices.

Frontier models fail quietly. Their degradation comes from corruption of existing content — values flipped, references rewritten, edits applied in the wrong place — producing documents that look intact at a glance but contain accumulated drift. The corruption mode is more dangerous than the deletion mode precisely because it preserves the surface signal of competence. The thing that looks like a successful workflow output is the thing that has silently drifted.

This matters for adoption. The "frontier models are reliable" intuition is built from short-interaction benchmarks where the corruption mechanism barely activates. At workflow scale — the regime where delegation is actually useful — the failure changes character, and the qualitative shift toward harder-to-detect failures means that improvements in raw capability can degrade overall workflow reliability if review effort is held constant.

The implication for delegated-AI design is that capability improvements at the frontier need to be paired with detection mechanisms that target corruption-style errors, not just deletion-style errors. Diff review, document-state checksums, and constraint validators become more important as models get better, not less.

Inquiring lines that read this note 70

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What compositional reasoning failures limit large language models despite scale? How do evaluation practices shape which failures stay visible? What mechanisms preserve shared understanding in evolving conversations? Can harness architecture and protocols provide agent reliability without model scaling? What training data selection strategies maximize generalization across difficulty levels? Why don't LLMs reliably translate capability into accurate outputs? How do capability benchmark scores systematically misrepresent true model abilities? What attack surfaces do reasoning traces and chains introduce? How should agent systems validate and persist generated code artifacts? How can infrastructure records verify actual agent behavior? Can validator consensus certify semantic correctness beyond agreement? What capability trade-offs arise from domain specialization through fine-tuning? Why do agents falsely report success on failed tasks? Can prompt-based context override biases that were embedded during pretraining? Can brute-force automated research substitute for iterative depth and human research intuition? How can we detect and prevent harm propagation through multi-agent delegation workflows? When do semantic similarity approaches miss structural retrieval failures? How does harness optimization generalize across different model architectures and domains? Do reasoning benchmarks predict model performance in long-horizon workflows? How do multi-agent LLM systems fail distinctly compared to single agents? How do social dynamics distort aggregated online ratings? Why do locally safe actions create system-level safety gaps? How can humans maintain meaningful oversight as AI systems become increasingly autonomous and complex? Do backend defenses obscure real attack effectiveness in reported metrics? What fundamental constraints limit how effectively agents can improve themselves? How do standardized protocols improve multi-agent coordination and reliability? Why do some clarifying approaches produce understanding while others just satisfy? Can we reliably detect when models game evaluations? Does AI assistance promote real skill development or substitute for independent learning? How effectively can language models perform reasoning, especially combined with symbolic methods? Is language model reasoning authentic and what causes models to reason? Do reasoning traces faithfully reflect actual model reasoning? How much do training data properties shape model reasoning? Do language models reason like humans or mimic surface patterns? Why do standard benchmarks fail to predict agent deployment success? How should designers communicate what AI systems truly are and can do?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
14 direct connections · 118 in 2-hop network ·medium cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

document degradation has a model-tier signature — weaker models delete content while frontier models corrupt it making frontier failures harder to detect