INQUIRING LINE

What does it actually look like when a company writes its know-how into a form AI can read and act on?

What does a machine-legible ontology look like in practice inside enterprises?

This explores what it looks like in practice when a company writes down its knowledge (its categories, rules, procedures and identifiers) in a structured form that AI systems can read and act on, rather than leaving it in people's heads or in loose prose.


This explores what a company's knowledge looks like once it is written down in a form AI systems can read and act on, beyond loose documents and people's heads. The corpus has no direct case studies of enterprise ontology projects. What it does have is a set of working examples of the pieces. Together they suggest that a machine-legible ontology in practice is less a single grand schema and more a stack of four concrete things: shared document formats, written-down rules, generated category systems, and identifiers that carry meaning.

The clearest example is procedure turned into document templates. MetaGPT gives agents the standard engineering documents a human team would use, such as requirements specs and design docs, instead of letting them chat Does structured artifact sharing outperform conversational coordination?. Agents pull what they need from a shared workspace, and coordination improves because the structure of the artifact carries the organization's logic. An industrial case study pushes the same idea further. Domain rules and design principles were built into an agent's scaffolding, the structured setup that wraps the model. Non-experts then produced work that experts rated as expert-level, a 206% quality gain that came from writing down tacit expertise rather than from a bigger model Can codified expertise let non-experts match specialist output?. In practice, then, the ontology often lives in the harness around the model: templates, rule sets, and checklists the agent must pass through.

Companies rarely have a clean taxonomy to start with, so the second piece is building one. TnT-LLM shows a practical route. An LLM reads messy text, proposes and refines a set of labels, labels the data, and then hands the work to cheap classifiers that run at scale Can LLMs efficiently generate taxonomies and label training data?. Identifiers matter too. Work on recommendation systems finds that neither bare ID numbers nor plain-text names work well alone. Combining ID, title and attributes gives you uniqueness, meaning, and something the model can reliably point to Can item identifiers balance uniqueness and semantic meaning?. That is a useful design lesson for any enterprise catalog of products, customers, or parts.

The less obvious lesson comes from mathematics. Efforts to translate math into formal, machine-checkable language find that formalizing one statement only works if you can borrow a whole prebuilt library of definitions and lemmas. The real work is the coherent web underneath Can autoformalization work on individual statements alone?. Enterprises face the same trap: tagging individual documents is easy, but the value comes from a connected set of definitions that hold together. Knowledge graphs give one way to make that web usable at run time. Small models that write their reasoning out as graph triples (subject–relation–object facts) solve much harder tasks, and every step can be inspected Can structuring reasoning as knowledge graphs help smaller models solve complex tasks?. Algorithmic control can then show each LLM call only the slice of structure it needs Can algorithms control LLM reasoning better than LLMs alone?.

The caution you might not expect: once AI edits these structured stores, the failure mode changes with model strength. Weaker models visibly delete content. Frontier models quietly corrupt it while keeping the document looking intact Does model capability change how documents degrade?. So a machine-legible ontology needs more than a schema. It also needs validation that checks meaning, not just format, because the better your models get, the harder their errors are to spot.


Sources 8 notes

Does structured artifact sharing outperform conversational coordination?

MetaGPT demonstrates that agents producing standardized engineering documents achieve superior coordination compared to conversational exchange. Active information pulling from shared environments eliminates noise and mirrors efficient human workplace infrastructure.

Can codified expertise let non-experts match specialist output?

An industrial case study embedding domain rules and design principles into an LLM agent's scaffolding achieved 206% output-quality improvement and expert-level ratings from non-experts, bypassing the need for specialist oversight. The capability gain came from externalizing tacit expertise into structured harness components, not from model scale.

Can LLMs efficiently generate taxonomies and label training data?

TnT-LLM automates text mining by using LLMs for open-ended reasoning to create and refine label taxonomies and generate training labels, then distilling these into lightweight classifiers for cost-effective deployment at scale.

Can item identifiers balance uniqueness and semantic meaning?

TransRec shows that combining numeric IDs, titles, and attributes into structured identifiers solves three problems simultaneously: distinctiveness from IDs, semantics from text, and generation grounding from structural constraints. Neither pure IDs nor pure text alone achieves all three.

Can autoformalization work on individual statements alone?

Real formalization requires theory-level work: even one theorem needs a coherent web of axioms, definitions, and lemmas. Statement-level approaches only succeed by borrowing from prebuilt libraries like Mathlib, hiding the actual complexity involved.

Show all 8 sources
Can structuring reasoning as knowledge graphs help smaller models solve complex tasks?

Knowledge Graph of Thoughts (KGoT) achieves 29% improvement on GAIA Level 3 tasks using GPT-4o mini by externalizing reasoning into iteratively constructed KG triples. The approach improves transparency, reduces bias, and enables quality control over reasoning steps.

Can algorithms control LLM reasoning better than LLMs alone?

LLM Programs embed LLMs within explicit algorithms that manage control flow and state, presenting only step-specific context to each LLM call. This information hiding addresses capability and context window limits while treating complex reasoning as modular, debuggable sub-tasks.

Does model capability change how documents degrade?

DELEGATE-52 shows weaker LLMs degrade documents through visible deletion, while frontier models degrade through subtle corruption that preserves surface integrity. This shift makes frontier failures harder to detect and potentially more dangerous at workflow scale.

Papers this line draws on 8

The research behind the notes this line reads — ranked by how closely each paper relates.