Can code serve as the operational substrate for agent reasoning?
Explores whether code functions not just as LLM output but as the executable medium through which agents reason, act, and verify progress. This reframing treats code as infrastructure rather than deliverable.
Most discussion of LLMs and code treats code as a product: the model writes a function, solves a competition problem, or patches a repository, and the code is the deliverable. The "code as agent harness" framing inverts this. In agentic systems, code is increasingly the operational substrate rather than the output — the medium through which an agent reasons (program-aided reasoning externalizes intermediate computation into executable form), acts (robotic and embodied agents run generated programs as policies), models its environment (codebases, execution traces, and tests represent state and dynamics), and verifies (runtime feedback confirms or refutes progress). What makes code uniquely suited to this role is that it is simultaneously executable, inspectable, and stateful: it can be run, read, and carried forward across steps.
This reframing connects threads that otherwise look separate — tool use, planning, memory, and verification all become facets of a single code-centered execution loop. The counterpoint is that not all agent reasoning reduces to code; natural-language deliberation and learned policies do real work that no program captures, and forcing everything into code can be a leaky abstraction. But where verification matters, code's executability gives agents a ground truth that prose lacks. This matters because it offers a unified lens for agent infrastructure: design the code substrate well and reasoning, action, and verification improve together.
Inquiring lines that read this note 75
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Should agents decouple planning from perception grounding for better performance?- Why do planning and grounding have opposing optimization requirements in agents?
- How should agents separate planning from perception grounding?
- Does the planning-grounding factoring principle apply to other agent tasks?
- How should the surrounding agent system be designed to ground actions in reality?
- How does credit assignment drive agents to write information into environments?
- Does encoding governance into runtime loops scale as deployment environments become more complex?
- Why do analysts prefer visible structured interfaces over hidden agent memory systems?
- Why do workflow abstractions fail in embodied agent environments?
- How do standardized artifacts improve coordination between writing agents?
- How do standardized artifacts reduce inter-agent communication failures?
- How do language agents implement prompts as executable computational graphs?
- Why do a-priori procedural specifications fail as environments change and interfaces evolve?
- What makes a service visible to autonomous agent systems?
- How should we measure context efficiency and verification cost in agents?
- Why do production AI agents deliberately stay simple and avoid frameworks?
- Can code-based reasoning replace natural language deliberation in agentic systems?
- What makes composable abstractions emerge under performance pressure in agent systems?
- When does forcing agent reasoning into code become a leaky abstraction?
- Does codifying domain rules into agent scaffolding work at library scale?
- What role do material artifacts play in solidifying AI relationships?
- What task characteristics determine whether humans or agents should handle work?
- Can deterministic function calls prevent agent failures better than protocol-mediated tool access?
- How do agents discover and construct new APIs from existing applications?
- What execution-layer design prevents agents from passively reacting to environments?
- How does protocol mediation affect determinism in agentic function calls?
- Why do production agents depend more on their surrounding pipeline than the model?
- Can open agent workflows be modeled as finite event lifecycles?
- Can API-first interaction replace traditional UI-based agent interfaces?
- Can specialized perception components replace end-to-end vision in GUI agents?
- Can automated evaluation replace human judgment in agent testing?
- What role does runtime feedback play in agent verification and progress confirmation?
- What other agent behaviors besides citations reveal reasoning quality?
- How do you verify agent code under incomplete feedback signals?
- Can algorithmic control flow over prompts simulate traditional programming languages?
- What makes language an effective parameterization for procedural knowledge?
- How does program-aided reasoning externalize computation into executable form?
- Can multi-agent debate prevent reasoning models from amplifying errors?
- What does collaborative computation mean when agents exchange and repair reasoning together?
- When does collaboration help versus harm in multi-agent reasoning?
- How should harness infrastructure validate code that agents generate themselves?
- When should agent-created code be promoted into permanent harness infrastructure?
- How do agents decide which created code should persist versus disappear?
- How should human oversight apply to persistent agent-authored code?
- Can one-off agent code be safely promoted to durable infrastructure?
- What makes persistent, shared code artifacts from agents hard to manage at scale?
- How do agent-created code artifacts become part of harness infrastructure?
- How do agents decide which created code deserves long-term persistence?
- How should agents decide which created code is worth persisting?
- Can disposable agent-authored code be distinguished from reusable infrastructure?
- How do execution traces represent state and dynamics in codebase modeling?
- At what interaction length does MCP's application-layer state code become unwieldy?
- What distinguishes communicative acts from operational actions in agentic LLMs?
- How do LLM-based agents develop shared abstractions through interaction?
- Why does forcing agents to trace function paths prevent unsupported claims?
- How do execution traces and tests represent agent environment state?
- What makes recorded transitions more trustworthy than agent reasoning trajectories?
- Can execution traces reveal unsupported claims in AI agent behavior?
Related concepts in this collection 2
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Should LLMs handle abstraction only in optimization?
What if LLMs worked exclusively on translating problems to formal constraints, while deterministic solvers handled the numeric work? Explores whether this division of labor could overcome LLM failures in iterative computation.
both treat emitting executable code as the locus of reliable reasoning rather than as a final answer
-
Can structured reasoning replace code execution for RL rewards?
Can semi-formal templates enable execution-free code verification reliable enough to train RL agents without running code? This matters because execution is expensive and slow in agent training loops.
explores the inspectable side of code as a reasoning medium even without execution
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Code as Agent Harness
- Agentic Code Reasoning
- RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- Emergent Hierarchical Reasoning In LLMs Through Reinforcement Learning
- BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure
- interwhen: A Generalizable Framework for Steering Reasoning Models with Test-time Verification
- Agentic Reasoning for Large Language Models
Original note title
code is not only llm output but an executable inspectable stateful medium through which agents reason act and verify