Does mapping your business into a formal model actually make an AI agent better at real work, or does other structure matter more?
Does ontology investment actually improve agent performance on business tasks?
This explores whether building a formal model of your business (an ontology: the defined entities, relationships and rules of your domain) makes AI agents better at real work. The corpus has no study that tests ontologies directly, so this answer draws on nearby research about what kinds of structure help agents.
This explores whether building a formal model of your business makes AI agents better at real work. The collection has no paper that runs this experiment, so it can't give a straight yes or no. It does have a lot on a closely related question: what kind of structure, given to an agent, actually changes its performance? The answer is more specific than "structure helps."
The strongest signal is that reliable agents get their reliability from structure built around the model, not from the model itself. The work on harness design argues that agents become dependable when memory, procedures and interaction rules move out of the model's head and into the system around it Where does agent reliability actually come from?. MetaGPT makes this concrete. Agents that pass standardized documents to each other, like specs, designs and interface definitions modeled on human operating procedures, coordinate better than agents that just talk Does structured artifact sharing outperform conversational coordination?. An ontology is one form of this kind of shared, agreed-upon structure. Seen that way, the evidence leans in its favor.
The surprise is what kind of structure does the work. A study of 8,135 agent trials found that when added "skills" helped, about two-thirds of the time it was because they steadied the agent's steps: what to do next, in what order. Only 4.5% of the gains came from supplying facts the agent lacked Do skills teach procedures or inject missing facts?. That matters for ontology spending. If your ontology mostly lists what things are (definitions, categories, attributes), it may be the part agents benefit from least. If it encodes how things move (which entity triggers which process, which state can follow which), it is closer to what the evidence says works. Related work suggests code may be a better home for that operational structure than descriptive schemas, because code can be run, inspected and checked Can code serve as the operational substrate for agent reasoning?.
There is also a counter-current. One line of work replaces hand-built routing and categorization with learned semantic matching. Agents are described by versioned capability vectors and found by similarity, with no manually maintained map Can semantic capability vectors replace manual agent routing?. This is the old argument between hand-built taxonomies and learned embeddings, now playing out in agent design. It suggests that some of what an ontology team would build by hand can be learned and kept current more cheaply.
Finally, "does it improve performance" may be the wrong measure. Evaluation research shows that identical task-success rates can hide large differences in efficiency, reliability and how ready a system is for deployment How should we measure agent system performance beyond task success?. An ontology's payoff might show up there instead: fewer wasted steps, easier checking, less drift between agents. A simple pass/fail benchmark would miss it. If you test this in your own organization, measure the agent's path to the answer, not just whether it got there.
Sources 6 notes
Research shows reliable LLM agents externalize three cognitive burdens—memory (state persistence), skills (procedural components), and protocols (structured interaction)—into a harness layer rather than relying on model scale alone. The harness unifies these externalities and eliminates the need for the model to solve the same problems repeatedly.
MetaGPT demonstrates that agents producing standardized engineering documents achieve superior coordination compared to conversational exchange. Active information pulling from shared environments eliminates noise and mirrors efficient human workplace infrastructure.
Analysis of 8,135 trials shows procedural anchoring accounts for 65.7% of skill cases versus 4.5% for knowledge injection. Skills fail when retrieved incorrectly, invoked out of context, or followed too rigidly.
Research shows code uniquely enables agent reasoning, action, and verification by being simultaneously executable, inspectable, and stateful. This unified code-centered loop improves reasoning and verification together compared to natural-language or prose-based approaches.
Versioned Capability Vectors embedded in HNSW indices couple semantic matching with policy and budget constraints, making capability discovery a first-class operation that scales sub-linearly as agent heterogeneity increases.
Show all 6 sources
Single task-success metrics obscure how agents achieve results across memory, context, and verification layers. Research shows identical success rates can mask enormous differences in efficiency, reliability, and deployment readiness—requiring harness-level benchmarks that measure trajectory, memory hygiene, and verification costs.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- Towards a Science of Scaling Agent Systems
- Demystifying Agent Skills: Why They Work-Until They Don't
- BenchShield: Formal Model-Backed Instrumentation for Reward Integrity in LLM-Agent Evaluation Infrastructure
- LatentSkill: From In-Context Textual Skills to In-Weight Latent Skills for LLM Agents
- MUSE-Autoskill: Self-Evolving Agents via Skill Creation, Memory, Management, and Evaluation
- RESOURCE2SKILL: Distilling Executable Agent Skills from Human-Created Multimodal Resources
- AgentCompass: A Unified Evaluation Infrastructure for Agent Capabilities
- From Model Scaling to System Scaling: Scaling the Harness in Agentic AI