Where does agent reliability actually come from?
Exploring whether LLM agent performance depends on larger models or on thoughtful system design choices like memory, skills, and protocols that shift cognitive work outside the model.
Drawing on Norman's concept of cognitive artifacts, this paper argues that the most consequential design choices in LLM agents are about externalization — relocating cognitive burdens from the model's internal computation into persistent, inspectable, reusable external structures. A shopping list doesn't expand memory; it changes recall into recognition. The same logic governs agent design.
Three dimensions of externalization address three recurrent mismatches:
Memory externalizes state across time. The context window is finite and session memory is weak. Memory systems transform recall into recognition — the agent retrieves past knowledge from a persistent store rather than regenerating it from weights. This solves the continuity problem.
Skills externalize procedural expertise. Long multi-step procedures are rederived rather than executed consistently. Skill systems transform generation into composition — the agent assembles behavior from pre-validated components rather than improvising each step. This solves the variance problem.
Protocols externalize interaction structure. Interactions with tools, services, and collaborators are brittle when left to free-form prompting. Protocols transform ad-hoc coordination into structured contracts (e.g., MCP). This solves the coordination problem.
The harness is not a fourth dimension — it is the engineering layer that hosts all three and provides orchestration logic, constraints, observability, and feedback loops. The progression is: weights → context → harness, paralleling the human history of cognitive externalization (speech → writing → printing → computation).
Critical system-level couplings:
- Memory expansion competes with skill loading for scarce context budget
- Protocol standardization can constrain how capabilities are packaged
- Skill execution generates traces that become memory; memory retrieval influences which skills and protocols are chosen
This reframes the question from "how capable is the model?" to "what burdens have been externalized so the model no longer has to solve them internally every time?" The base model may remain unchanged; what changes is the representation of the task.
This connects to Why do production AI agents stay deliberately simple? — the externalization framework explains why custom harnesses outperform: they externalize the right cognitive burdens for their specific domain. It also extends When should human-agent systems ask for human help? — Magentic-UI's mechanisms (co-planning, action guards, memory) are specific instances of the three externalization dimensions.
The "From Model Scaling to System Scaling" paper sharpens this into an explicit framing: model scaling (bigger models, more data, higher benchmark scores) versus system scaling (designing the auditable, persistent, modular, verifiable architecture around the model). It treats the harness as a first-class object of design, evaluation, and optimization, decomposing it into a foundation model, memory substrate, context constructor, skill-routing layer, orchestration loop, and verification-and-governance layer — a finer-grained partition of the same memory/skills/protocols externalization. Its central demonstration is that comparable models projected onto different harnesses (Claude Code, OpenClaw, and the released CheetahClaws reference harness) produce qualitatively different agents, making the harness "now a primary source of practical capability." This is direct evidence for the claim that reliability comes from the surrounding system, not from a larger model alone.
Inquiring lines that read this note 348
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What makes agent memory systems durable and reusable across sessions?- Can persistent memory and identity files alone create genuine agent socialization?
- Does state persistence in AI systems create the same temporal presence as human waiting?
- How should GUI agents remember patterns across different software environments?
- Can environmental scaffolding replace internal memory scaling in agent design?
- Could a single agent system switch memory granularity between tasks?
- What memory and planning capabilities do AI companions need for evolving user needs?
- Why do memory and feedback loops matter more than model size for agent reliability?
- Can episodic memory of UI traces improve open-world agent adaptation?
- Can state-indexed memory retrieval breadth predict gains in web agent robustness?
- How does PRAXIS differ architecturally from Agent Workflow Memory and causal rule learning?
- Which memory components trigger context-length problems in agents?
- Can multimodal agents use entity-centric graphs within this three-axis framework?
- Can pruning policies alone solve working memory bloat in agents?
- How does workflow abstraction compare to state-indexed procedural memory for web agents?
- Can agent-controlled memory management outperform fixed consolidation schedules?
- Does workflow-level memory or state-action memory better capture reusable agent knowledge?
- Can AI models retain knowledge across changing environments without catastrophic forgetting?
- How do planning and memory compress agentic system costs?
- What distinguishes working memory from strategic memory in agent task execution?
- How does durable memory quality shape agent performance over time?
- How do memory tools and planning each contribute to agent efficiency?
- How does external context control compare to agents managing their own state internally?
- How do memory hygiene and context efficiency trade off in deployed agents?
- What causes multi-turn agent failures: weak memory control or missing knowledge?
- How does structured environment-side state reduce multi-turn agent failure better than transcript replay?
- Can workflow memory compound reusable skills into measurable success improvements?
- Why do persistent AI systems require fundamentally different design than ad-hoc supporters?
- Why do analysts prefer visible structured interfaces over hidden agent memory systems?
- How do workflow and function memories contribute differently in agent learning?
- Should memory type shape what kind of agent responses work best?
- How do modern agents separate fast non-parametric updates from slow weight learning?
- How does the agentic layer amplify individual agent failure modes?
- Do architectural changes or training fixes better prevent agreement failures?
- Why do homogeneous multi-agent systems fail similarly to self-revision?
- Do multi-agent systems justify their token costs with genuine quality gains?
- Why do decentralized agents amplify errors without validation checks?
- How does collaboration topology choice affect error amplification in multi-agent systems?
- Which failure mode most limits current multi-agent performance?
- Where should the trust boundary sit in multi-agent planning systems?
- What degradation patterns emerge as relay length increases in delegated tasks?
- What makes observation and intervention placement different across agent pipelines?
- What does error recovery look like across different agent architectures?
- Can correct verdicts hide failures in agent coordination steps?
- Where should the trust boundary sit in multi-agent planner systems?
- Does adding capability without improving detection reduce overall system reliability?
- How prevalent is misaligned behavior in dense multi-agent interaction settings?
- Are deployed agents typically settled about their objectives by design?
- Why do multi-agent failures arise through interactions local checks miss?
- Why do comparable metrics matter across different multi-agent system designs?
- Can multi-agent architecture isolation reveal which design choices matter most for safety?
- What does collaborative computation mean when agents exchange and repair reasoning together?
- Which eleven failure modes emerge from agentic layers in realistic deployment?
- How often do multi-agent systems fail from provider refusals versus agent errors?
- Can procedural instructions and platform checks recover performance lost by multi-agent teams?
- How do multi-agent LLM systems fail at coordination and role consistency?
- Can parallel agents or complementary mechanisms replace single-human interrogation of LLMs?
- Why do LLM agents make promises without executing them?
- Why do LLM agents fail where game-theoretic bots succeed?
- What makes LLM agents default to passive helpfulness without curiosity rewards?
- Do parallel LLM workers coordinate emergently without predefined collaboration rules?
- Can multi-agent LLM systems overcome diversity collapse through structured disagreement?
- Do agent frameworks adequately compensate for LLM conversational passivity?
- Can LLMs coordinate with humans better using different model architectures?
- How do shared KV caches enable emergent coordination between LLM agents?
- Why do LLM agents struggle with protocol discipline in distributed settings?
- Do multi-agent language model teams fail the same way individual reasoning does?
- What distinguishes communicative acts from operational actions in agentic LLMs?
- What role should reasoning agents play in validating multi-LLM ensemble outputs?
- How do specialized agent roles improve consistency in long-form writing?
- Do multi-agent LLM systems scale better than centralized hierarchies?
- Do multi-agent LLM systems fail in measurably different ways than single agents?
- How does agent reliability emerge from memory and protocols instead of model scale?
- Does multi-agent interaction amplify existing failures or create new ones?
- How do LLM-based agents develop shared abstractions through interaction?
- How do postmortem convention-setting stages enable language evolution in agents?
- Why did older robot scientists log auditable provenance while LLM agents inherited only generation capability?
- How do LLM user simulators track and maintain consistent goal states across multi-turn interactions?
- Do emotion-driven actions in agent simulators capture genuine belief revision or just reactive behavior?
- Why do longer forecasting horizons degrade LLM accuracy in role-play?
- What distinguishes a neutral simulator from an agent with its own agency?
- Does adjusting steered mechanisms make LLM agents match human behavior more closely?
- Why do planning and grounding have opposing optimization requirements in agents?
- How should agents separate planning from perception grounding?
- Does the planning-grounding factoring principle apply to other agent tasks?
- How do planning and grounding have opposing optimization requirements in agents?
- Do GUI agents need harness-level splits between planning and grounding?
- Should user simulators be trained via RL like agents or decomposed into trackable state components?
- What makes software engineering environments better suited for RL than other interactive domains?
- How does credit assignment drive agents to write information into environments?
- Why do weak belief tracking and conservative actions trap agents in low-information states?
- Can applicability conditions be preserved automatically when agents reflect on trials?
- How do human-agent systems incorporate diverse feedback into model behavior?
- How do agents differ in caution versus persistence across low-information scenarios?
- How does effective feedback retention govern long-horizon agent reliability?
- Why does persistence in the feedback loop predict agent success better than initial solution quality?
- Can agents improve reliably without an external standard?
- Why does human interaction remain the hardest failure mode for agents?
- Why do agents report success when they have actually failed at tasks?
- Can agent success reports serve as reliable oversight signals in real deployment?
- How much autonomy can agents safely exercise before failing?
- What tasks do AI agents still fail at most often?
- Why do completion-mode strengths not transfer to agentic settings?
- How do mode-specific failures differ between completion and agent benchmarks?
- Why do agents report success when actions actually fail?
- How do agents learn to report success on actions that actually failed?
- Where does agent reliability come from if not better tools?
- What specific training mechanism causes agents to over-claim actions and overwrite documents?
- Why do AI agents fail at verification but succeed at generation?
- What makes idle window detection valuable for continuous agent improvement?
- Which failure modes dominate in autonomous research agents?
- When should agents stop recursing to optimize success versus cost?
- How should safety systems catch confident failures from agents that report success on unsafe actions?
- How does completion bias in agents differ from other epistemic failure modes?
- How do agents decide when to stop and reflect on failure?
- What hidden signals in agent logs reveal about frontier capability beyond pass-fail outcomes?
- Why do AI agents struggle with novel experiments but excel at routine tasks?
- Can stopping rules extracted from past failures improve agent reliability without retraining?
- How does poor belief tracking cause agents to keep acting past the point of usefulness?
- Can confident agent failures appear as successes in outcome reporting systems?
- How often do agents report success when their actions actually failed?
- Why do autonomous agents report success on failed actions?
- What causes the gap between agent reasoning and agent action?
- Why do agents report success when their actions actually fail?
- How do agent accuracy and error recovery affect delegation time?
- Why do autonomous AI agents fail at real workplace tasks?
- Why is complex UI navigation the hardest agent failure mode?
- Can smaller models trained for execution handle the failure modes that stop current agents?
- Why do long-horizon agents fail when their models can solve individual steps?
- How do domain experts recover from agent errors differently than novices?
- Why do novice users abandon troubled agent sessions three times more often?
- Why do private knowledge domains remain the hardest failure mode for agents?
- What makes users willing to relinquish control to an agent?
- How should AI systems model human resource constraints and expertise levels?
- What task characteristics determine whether humans or agents should handle work?
- How does machine agency spectrum explain tool design mismatches with user behavior?
- How should humans and AI agents share decision-making authority?
- What conditions let users configure agents to match their priorities?
- Does the planning-execution split between humans and agents depend on policy?
- What autonomy levels do workers prefer compared to actual agent deployments?
- Why do workflow abstractions fail in embodied agent environments?
- Why do rigid orchestration frameworks fail where generative environment specifications succeed?
- Can deterministic function calls prevent agent failures better than protocol-mediated tool access?
- How do standardized artifacts prevent autonomous agent failure modes?
- What role does standardization play in multi-agent system ecosystems?
- How do standardized artifacts improve coordination between writing agents?
- How do standardized artifacts reduce inter-agent communication failures?
- How should human oversight apply to persistent agent-authored code?
- Can one-off agent code be safely promoted to durable infrastructure?
- Should new agent protocols replace existing ones or layer on top of them?
- What would unified agent-to-agent and agent-to-tool protocols actually look like?
- How do agents decide which created code deserves long-term persistence?
- Can open agent workflows be modeled as finite event lifecycles?
- What domain properties determine whether causal rules transfer to new agents?
- When should you optimize agent behavior versus tool performance separately?
- Can agentic reasoning outperform rigid rule-based systems for skill refinement?
- What happens when agents interact with environments and learn from their own mistakes?
- How much does agent performance depend on demonstration quantity versus curation quality?
- Can agents improve from deployment signals without explicit human annotation?
- Should agent capability be optimized separately from general capability?
- Can agentic AI tools deliver productivity gains on learning tasks differently?
- How do agent capabilities change across 25 relay rounds of interaction?
- How do agents automatically generate suitable learning tasks based on current capability?
- What lifecycle management prevents in-loop skill creation from bloating an agent?
- How do fast and slow timescales enable continual agent adaptation?
- What properties of agent systems only become visible across multiple sessions?
- Can context management policies transfer across agents of similar capability levels?
- Can simulation fidelity limit what agents learn from trained world models?
- Can agent skills move from prompts to trainable parameters?
- Can agent-authored skill libraries compound autonomy gains over time?
- Where does an agent's risk come from across its components and sequence?
- How much does external context management transfer across similar capability agents?
- How can agent data flywheels improve task quality iteratively?
- What makes an agent mechanism reusable versus benchmark-specific?
- Can agents acquire new skills online when offline skill coverage runs out?
- How do parametric and non-parametric updates differ in agents?
- How do complexity, diversity, and real-world fidelity interact in agent training?
- Do agents actually convert raw experience into better behavior automatically?
- Does ontology investment actually improve agent performance on business tasks?
- Can the scaling law for discovery extend beyond architectures to agentic systems?
- How do cognitive stimulation and process losses interact in group AI systems?
- What distinguishes collective evolution from vertical self-improvement in agent systems?
- What accounts for performance drops in multi-turn agent interactions?
- How do multi-agent systems improve on single frontier models?
- Does upgrading model capability improve token efficiency in agentic systems?
- Does parallel task structure determine optimal multi-agent architecture?
- Can cognitive diversity overcome expertise gaps in agent teams?
- Can cognitive diversity compensate for lack of expertise in agent teams?
- Why do 85 percent of production agents avoid third-party frameworks?
- What ecosystem conditions make agent attention markets viable?
- Which layer of agent systems creates the largest capability gains in practice?
- How should proportionality constraints be implemented in agentic systems?
- Why do production AI agents deliberately stay simple and avoid frameworks?
- How do externalizing cognitive work and coordination infrastructure relate to agent reliability?
- Why does capability discovery become the bottleneck in large agent systems?
- Can code-based reasoning replace natural language deliberation in agentic systems?
- What makes composable abstractions emerge under performance pressure in agent systems?
- How do capability vectors enable discovery in multi-agent systems?
- How does deterministic feature engineering increase information for computationally bounded agents?
- Can we design efficient agents by targeting constraints directly?
- Can multi-agent teams solve problems better than single models thinking longer?
- How will the agent economy reshape compute infrastructure design?
- How do perception and execution gaps limit current AI agent performance?
- When does forcing agent reasoning into code become a leaky abstraction?
- How does multi-agent reasoning scale compared to single-model approaches?
- What structural features drive instrumental convergence across different agent goals?
- Why does sycophantic relay propagate planning-time bias through agent pipelines?
- Do specialized agents outperform single agents with better orchestration?
- Do recursive subagents reduce single-model context pressure?
- Do agent-created languages improve or degrade performance on their original tasks?
- Is the coupled human-agent environment the right unit for evaluation?
- Can AI agents benefit from relational traits like consistency and prosociality?
- Can prompt engineering fully prevent role flipping in LLM agents?
- How do external prompt artifacts improve agent behavior compared to inline instructions?
- What distinguishes strategic fabrication from accidental hallucination in research agents?
- Does transparency in policy language improve agent trustworthiness over time?
- What other agent behaviors besides citations reveal reasoning quality?
- Do agents systematically misreport their own capabilities and tool access?
- What does loss of control language assume about agent cognition and intent?
- What distinguishes domain-specific failure modes from general model limitations?
- Can models optimized for solo capability support productive human collaboration?
- Why do high-level design guidelines fail to capture real-world deployment nuance?
- Which model capabilities actually matter for sustained workflow delegation?
- Can architectural changes reorder when uncertainty and empowerment signals influence decisions?
- How do agentic systems recover when specialized models operate outside their scope?
- What governance and safety measurements matter for deployed agent environments?
- What specific bookkeeping tasks can environments maintain more reliably than policies?
- Why has agent research prioritized policy over world model development?
- Can a single manager policy work across vastly different agent architectures?
- What counts as a mature governance model for agentic AI systems?
- Do agents prefer raw experience over condensed summaries of past actions?
- Can agents compress long trajectories without losing critical decision context?
- When does memory consolidation help agents instead of hurting performance?
- Why do continuously consolidated agent memories eventually degrade below no-memory baseline?
- Why do agents systematically underuse condensed experience in skill documents?
- What separates artifact recall from persistent memory commitment in agents?
- Why do agents ignore condensed experience in favor of raw data?
- How does indiscriminate memory injection cause multi-turn agent failures?
- How does bounded committed state prevent multi-turn agent failures better than transcript replay?
- Does selective history retrieval outperform full context inclusion in agent reasoning?
- Does reducing interaction history cost agents performance on their tasks?
- Why do agents systematically ignore condensed experience in their skill documents?
- How should we evaluate agent memory if it folds into model computation instead of separate stages?
- Why do agents ignore condensed experience even when it is the only evidence available?
- Why do AI agents default to passivity when deferral timing is unclear?
- How do agents decide when to abstain from contributing?
- How should the surrounding agent system be designed to ground actions in reality?
- What execution-layer design prevents agents from passively reacting to environments?
- Why do production agents depend more on their surrounding pipeline than the model?
- What components of agent scaffolding most impact domain-specific output quality?
- Why does externalized state beat parameter scaling for agent reliability?
- How does externalizing reasoning into harness artifacts improve agent reliability?
- What makes agent-initiated artifacts the underexplored frontier in harness engineering?
- How do different harness designs produce different agent behaviors from the same model?
- Which harness dimensions most directly predict agent system reliability?
- How much realized agent capability comes from the harness versus the model?
- What makes some model capabilities reliable while others remain brittle?
- How does the LLM Fallacy differ from automation bias and cognitive offloading?
- What makes task alignment more fragile than underlying knowledge retention?
- What unique perspective do designers bring to LLM adaptation that engineers might miss?
- How much does omniscient evaluation overstate real-world simulation fidelity?
- Can test environments reliably predict how models behave in actual deployment?
- Should GUI agents use intermediate structured representations instead of raw pixels?
- Can screen perception be effectively decoupled from planning in GUI agents?
- What capability threshold do agents need to self-organize effectively?
- Why do current evaluation metrics fail to catch reasoning failures in persona agents?
- Which ecosystem conditions matter most for agent deployment success?
- How should we measure context efficiency and verification cost in agents?
- How do evaluation methods differ for single versus multi-agent systems?
- How should benchmarks measure agent efficiency across all three cost dimensions?
- Can single benchmarks predict whether an agent will work in the real world?
- Should artifact-level benchmarks replace token counts for agent evaluation?
- Does single-capability ranking guarantee agent failure in production deployment?
- Which agent architectures consistently outperform base models on hard prediction questions?
- Do trajectory quality metrics predict agent safety and user trust?
- Can single-axis benchmarks measure across all three agent capability layers?
- What agent evaluation dimensions beyond task success does a single number hide?
- Which interaction artifacts matter most for reliable agent evaluation?
- Should agent evaluation include trajectory quality beyond final success?
- Does a single benchmark score systematically misrepresent multi-axis agent capability?
- What dimensions beyond task success matter for evaluating long-horizon agent trajectories?
- Do automated benchmarks systematically distort what long-horizon agent capability actually looks like?
- What makes single-axis agent benchmarks unreliable for predicting deployment readiness?
- How do agent capability axes misalign with what users actually value?
- How do agent benchmarks misrepresent real-world deployment readiness?
- Why is the coupled human-agent environment the right unit of evaluation?
- Does episode-level cost become the decisive factor when comparing AI agents in production?
- Should agent evaluation include trajectory quality and memory hygiene alongside task success?
- Can a single leaderboard score capture multi-dimensional differences in agent performance?
- Can a single agent benchmark score accurately represent deployment readiness?
- Why do single-axis benchmarks fail to measure deployment-ready agent capability?
- How do single-axis benchmarks misrepresent AI agent readiness for deployment?
- Does agent capability separate into independent axes like performance and integrity?
- Does source bias affect real deployed agents or only benchmark environments?
- Why do high-scoring agents default to known techniques rather than novel solutions?
- Why do multi-agent systems converge without genuine deliberation?
- Why do multi-agent LLM systems converge prematurely without genuine deliberation or probing?
- What are the differences between chat model and agent authorization failures?
- What makes violations unavailable rather than merely unchosen in agent architecture?
- Can agents rationalize rule violations by reframing them as repairs?
- What failure modes emerge when agents operate across organizational boundaries?
- Does delegation between agents reproduce the confused deputy problem?
- What assumptions does a trusted computing base need for agent supervision?
- How do agents decide when to pause and reflect on their strategy?
- How do goal and environment choices mediate AI agent risk pathways?
- What role does runtime feedback play in agent verification and progress confirmation?
- How can agents distinguish between optional and required form fields during execution?
- How do you verify agent code under incomplete feedback signals?
- What permission models govern code execution within agent skills?
- When do agents benefit most from reusable workflow routines?
- Do agent improvements discovered on code tasks transfer to non-coding domains as well?
- How can agents detect missing information before attempting to solve problems?
- Why does continuous agent inference differ from human user inference?
- What trust signals do agents lack that humans use to assess credibility?
- What makes users trust an AI agent's proposed plan?
- Does codifying expertise into AI agents drive faster labor substitution?
- Can persistent agentic workflows predict labor displacement better than task-level exposure?
- Do frontier AI models fail in ways that preserve the appearance of competence?
- Can self-reported AI reliability metrics hide confounding factors like task complexity?
- Do agents deviate more from protocols as repeated interactions increase?
- Does restricting interaction history visibility reduce misaligned communication in agent markets?
- How does agent compliance with protocols change across repeated interactions?
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Externalization in LLM Agents: A Unified Review of Memory, Skills, Protocols and Harness Engineering
- LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents
- RRSI: Regularized Recursive Self-Improvement of Agent Harnesses
- Useful Memories Become Faulty When Continuously Updated by LLMs
- Bilevel Coordinated Reflection: A Game-Theoretic Approach to Multi-Agent LLM Systems
- Harness Updating Is Not Harness Benefit: Disentangling Evolution Capabilities in Self-Evolving LLM Agents
- Demystifying Agent Skills: Why They Work-Until They Don't
- Rethinking the Evaluation of Harness Evolution for Agents
Original note title
agent reliability comes from externalizing cognitive burdens into memory skills and protocols not from larger models — the harness is the unification layer