Line of inquiry
Inquiring lines›How can multi-agent systems achiev…›What conditions allow multi-agent…›this line of inquiry
How much of agent capability comes from harness versus the model itself?
A broader line of inquiry — a family of 19 specific questions the research asks around this. Follow one into its inquiring-line page, or move sideways to a related line below.
Questions in this line of inquiry 19
Specific inquiring lines the field asks around this — ordered from the most general framing down to the most specific angle.
- How does externalizing reasoning into harness artifacts improve agent reliability?
- How do agentic systems hide harness failures from benchmarks?
- What components of agent scaffolding most impact domain-specific output quality?
- How do different harness designs produce different agent behaviors from the same model?
- How much realized agent capability comes from the harness versus the model?
- Why does the harness layer accumulate distributed behaviors over time?
- Why does externalized state beat parameter scaling for agent reliability?
- What makes behavior localization the bottleneck in agent harness evolution?
- How do agent-created code artifacts become part of harness infrastructure?
- What makes harnesses more tangled than other types of agent code?
- Which harness dimensions most directly predict agent system reliability?
- How should the surrounding agent system be designed to ground actions in reality?
- What makes agent-initiated artifacts the underexplored frontier in harness engineering?
- Why do production agents depend more on their surrounding pipeline than the model?
- Why do persistent, resynchronized artifacts compound harness capability gains?
- What execution-layer design prevents agents from passively reacting to environments?
- What safety relations does a domain supply that a harness must capture?
- How can harnesses externalize bookkeeping so models focus on semantic judgment?
- What makes durable code artifacts more valuable than per-task harness patches?