INQUIRING LINE

Should AI research agents remember their work in a never-edited logbook or a running summary that keeps getting rewritten?

How do append-only Git records compare to continuously updated world models for research continuity?

This explores two ways a long-running research effort can remember what it has done: keeping a growing log that is never edited (an append-only Git history), or keeping one running summary of 'what we know' that gets rewritten as work goes on. The question is which one keeps research coherent across many sessions and agents.


This explores how research agents keep track of their work across sessions: by keeping an append-only history where old entries are never overwritten, or by keeping a running summary of current understanding that is rewritten as they go. The corpus has a strong case study of the first approach. It doesn't have a direct head-to-head comparison, but several notes show what goes wrong when you rely on the second. In Can decentralized agents coordinate research without a central planner?, thirteen language-model workers with no central planner shared a Git history for 12 days. They produced 1,703 contributions and closed 62% of the gap to a trained baseline. The history did the job a coordinator would normally do. Each new session could see what had been tried, what it came from, and what it produced, without rebuilding that context from scratch.

The main risk of a running summary is not obvious. Every rewrite is a chance to lose information. The document-relay studies show how big that risk is. When frontier models repeatedly edit a document over a long workflow, they corrupt about 25% of the content, and the damage slows down but never stops Do frontier LLMs silently corrupt documents in long workflows?. The type of damage also changes as models get stronger. Weaker models visibly delete content, while frontier models change it quietly in ways that still look fine on the surface Does model capability change how documents degrade?. A summary that an LLM keeps rewriting is exactly this kind of relay. An append-only log avoids the problem because nothing already written ever passes back through the model to be edited.

A second advantage is that a log keeps the failures. Can research papers preserve the experiments that failed? argues that published papers act as 'lossy compilers': they keep the story that worked and drop the dead ends, the reasoning behind decisions, and implementation details. A running summary tends to compress the same way, since its job is to state what is currently believed. A Git history keeps the rejected branches automatically. For anyone continuing the work, those branches are often the most useful part, because they show what not to try again.

Running summaries still have a real strength, which is orientation. A raw log is hard to read, and that has to be handled somewhere. Can building a document map first improve retrieval over long texts? shows that building an overall map of a document first helps retrieval connect evidence that is far apart, which a pile of fragments can't do. Can orchestration layers make coding agents more auditable? wraps coding agents in persistent state objects and reports a process trail that can be traced and recovered. That suggests a hybrid design: the log is the source of truth, and summaries are derived views of it. A summary that drifts can then be thrown away and rebuilt from the log.

This matches a broader principle in Can separating judgment from verification improve research paper reliability?: keep the model's judgment separate from records that are fixed and can be checked. An append-only log puts the record outside the model's reach, and the summary is just one interpretation of it. The useful takeaway is that the real question isn't log versus summary. It's whether your memory of the research can be quietly rewritten by the same system that is doing the research.


Sources 7 notes

Can decentralized agents coordinate research without a central planner?

Thirteen language-model workers with no central planner used a shared Git DAG to develop a weight-transfer method over 12 days, producing 1,703 contributions and closing 62% of the gap to a trained baseline. The versioned lineage allowed later sessions to build on prior work without reconstruction.

Do frontier LLMs silently corrupt documents in long workflows?

Even the strongest models (Gemini 3.1 Pro, Claude 4.6 Opus, GPT 5.4) degrade documents by ~25% over long relay workflows across 52 domains. Degradation decelerates but never plateaus, and errors compound silently, remaining undetected in spot-checked outputs.

Does model capability change how documents degrade?

DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.

Can research papers preserve the experiments that failed?

Publishing imposes a Storytelling Tax (erasing process, failed branches, tacit reasoning) and Engineering Tax (omitting implementation specs). Agent-Native Research Artifacts address both by packaging logic, executable code, exploration graphs of failures, and evidence grounding—treating rejected branches as publishable deliverables rather than editorial casualties.

Can building a document map first improve retrieval over long texts?

MiA-RAG inverts standard RAG by summarizing documents first, then conditioning retrieval on that global view. This approach recovers discourse structure that bag-of-chunks retrieval destroys, making scattered evidence findable by their document role rather than surface similarity alone.

Show all 7 sources
Can orchestration layers make coding agents more auditable?

Dr. Claw wraps existing coding agents in persistent state objects and skill libraries, reporting higher research completeness and a traceable, recoverable process trail while keeping the underlying executor unchanged.

Can separating judgment from verification improve research paper reliability?

Spark-to-Paper architects paper generation as composable skills that isolate model judgment from executable, verifiable operations and require evidence specification before results are observed, reducing dependence on model correctness for consistency.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.