SYNTHESIS NOTE
Topics›Knowledge After the Web›this note

Why do LLMs struggle with organizing long-form non-fiction?

Nathan Lambert's textbook writing experience suggests LLMs fail at integrating knowledge across chapters despite handling individual sentences well. The question is whether this reflects a fundamental limitation in how models compress and organize information.

Synthesis note · 2026-10-09 · sourced from Knowledge After the Web

Nathan Lambert, who just finished writing a post-training textbook (Reinforcement Learning from Human Feedback, Manning), argues that LLM progress on long-form non-fiction writing has stalled even as models have gone "from okay to superhuman" at coding and math. He writes that "the models seem genuinely horrible at long-form technical writing" — they can get a sentence or equation right, but asked to write "an entire chapter it'll be a mix of sprinkled with confusing wording, muddled in its organization, and generally a bit off." He used AI extensively for copyediting, LaTeX/TikZ diagrams, and syncing Markdown and LaTeX versions of the book, but says "less than 1%" of the book's explanatory sentences came from a model, included only where he, "as a true expert," judged them as what the reader needed.

Lambert's explanation is that "organizing knowledge is a compression" needed to produce insight, and that today's models do the opposite: they "increase entropy in long-form non-fiction writing." The failure mode he names is compounding error — models can check "every unit of content," a sentence, equation, or figure, but "don't do a good job revisiting components and stringing them together" across many additions, producing what he calls "irreducible compounding errors." He contrasts this with math and code, where he credits RLVR as "a truly magical solution" to the same compounding-error problem — implying long-form writing lacks an equivalent verifiable training signal.

This dovetails with Can LLMs generate more novel ideas than human experts?: Lambert's unit-level competence paired with integration failure is the same generation/evaluation split described there, applied to prose instead of research ideas. It also echoes Do LLM improvements reflect reasoning gains or corpus shifts? in locating the real source of quality in human judgment rather than model reasoning — Lambert says his own "understanding, in the form of intuition, taste, instinct" is what stayed valuable, and that using AI for the writing itself "takes away from that progression." By contrast, where Does polished AI output trick audiences into trusting it? warns that polished AI output can pass as expert judgment to an unwitting audience, Lambert describes the opposite outcome under expert supervision: because he kept "a very close eye" on the few AI sentences he kept, style never substituted for his own judgment.

This is one author's working experience with a small number of named models (GPT-4.5, Kimi K2, Claude) on one book, not a controlled study of writing quality, and Lambert himself is extrapolating a two-to-five-year forecast from it. It doesn't establish why long-form writing specifically lacks a tractable training signal, only his guess that it is "challenging and lacks good training data." The implication he draws, cautiously, is that if models can't yet organize and compellingly present already-established science, claims about them autonomously solving open scientific problems should be treated with skepticism until this narrower, prerequisite skill improves.

Inquiring lines that read this note 1

This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.

What limits language model accuracy in evaluating ideas?

Related concepts in this collection 4

This note in its neighbourhood — explore the map, then jump to a related concept in the list below.

Concept map
12 direct connections · 113 in 2-hop network ·dense cluster Open in graph ↗

Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph

your link semantically near linked from elsewhere

Related papers in this collection 8

Papers most semantically related to this note, ranked by cosine similarity in the embedding space.

Original note title

Lambert argues LLM writing ability has stalled on long-form non-fiction because organizing knowledge is a compression the models cannot yet do