Theme of inquiry
What conditions allow multi-agent systems to coordinate and reason well?
A question within its area, explored through 12 lines of inquiry below — each a family of specific questions the research asks.
73 specific questions
- Does a single benchmark score systematically misrepresent multi-axis agent capability?
- How do agent capability axes misalign with what users actually value?
- Can a single capability score hide an agent's tendency to game evaluations?
- How do agent benchmarks misrepresent real-world deployment readiness?
- What makes some agent benchmarks measure interaction quality better than others?
- Should agent evaluation include trajectory quality beyond final success?
- How do single-axis benchmarks misrepresent AI agent readiness for deployment?
83 specific questions
- Why do different agent memory architectures make incompatible granularity claims?
- Could a single agent system switch memory granularity between tasks?
- Can agent-controlled memory management outperform fixed consolidation schedules?
- How does durable memory quality shape agent performance over time?
- Does workflow-level memory or state-action memory better capture reusable agent knowledge?
- Should memory tools and prompts be governed structurally rather than patched?
- What makes agent-scale memory persist beyond individual sessions or people?
58 specific questions
- Can multi-agent debate prevent the confident convergence on wrong answers?
- What mechanisms drive silent agreement in multi-agent reasoning systems?
- Why do multi-agent systems converge on wrong answers without debate safeguards?
- Can silent agreement be prevented in multi-agent reasoning systems?
- Why does premature consensus form in multi-agent reasoning systems?
- Why does premature consensus form in multi-agent reasoning without genuine deliberation?
- What causes silent agreement in multi-agent reasoning systems?
54 specific questions
- Can agents learn to distinguish helpful from misleading interventions?
- Should feedback channels be excluded from the reward path in agent evaluations?
- Does reasoning ability help agents learn from feedback faster?
- Why do agents fail to internalize value from informative observations?
- What makes an agent notice that reward beats compliance?
- Can an agent's internal probabilities serve as value signals across domains?
- Can a reward-seeking agent be distinguished from one pursuing intended behavior?
88 specific questions
- Can correct verdicts hide failures in agent coordination steps?
- Which failure mode most limits current multi-agent performance?
- What distinguishes task failure from communication breakdown in multi-agent systems?
- How often do multi-agent systems fail from provider refusals versus agent errors?
- How do agreement-detection agents improve distributed coordination outcomes?
- How does distributed coordination fail as agent networks scale?
- Can procedural instructions and platform checks recover performance lost by multi-agent teams?
83 specific questions
- Do learned workflows transfer between different agents with minimal accuracy loss?
- Can skill repositories evolve toward execution-oriented refinement over time?
- Can individual skills improve through reuse and accumulate experience across tasks?
- Can agent skills move from prompts to trainable parameters?
- Can agent-authored skill libraries compound autonomy gains over time?
- Can agentic reasoning outperform rigid rule-based systems for skill refinement?
- Can agents learn beyond the boundaries of their training data curators?
60 specific questions
- Why do agents report success when actions actually fail?
- Why do agents report success when their actions actually fail?
- How often do agents report success when their actions actually failed?
- Why do agents report success when they have actually failed at tasks?
- How do agents learn to report success on actions that actually failed?
- Why do autonomous agents report success on failed actions?
- Why do agents claim completion when their outputs remain incomplete?
49 specific questions
- How do multi-agent LLM systems fail at coordination and role consistency?
- Do multi-agent language model teams fail the same way individual reasoning does?
- Do multi-agent LLM systems fail in measurably different ways than single agents?
- Does multi-agent interaction amplify existing failures or create new ones?
- How do LLM-based agents develop shared abstractions through interaction?
- What specific conditions allow language evolution in multi-agent LLM systems?
- Why do LLM agents struggle with protocol discipline in distributed settings?
98 specific questions
- How do single-agent capabilities affect the trade-off between coordinator and team architectures?
- Do single-agent systems outperform multi-agent coordination as model capabilities grow?
- Do agents inform neighbors when adopting strategies in their reasoning?
- How do multi-agent systems improve on single frontier models?
- When does multi-agent scaling actually outperform static ensembles?
- Can cooperative AI systems make meaningful decisions without a stable self?
- What accounts for performance drops in multi-turn agent interactions?
17 specific questions
- How do planning and grounding have opposing optimization requirements in agents?
- Why do planning and grounding have opposing optimization requirements in agents?
- Does the planning-grounding factoring principle apply to other agent tasks?
- Do GUI agents need harness-level splits between planning and grounding?
- What distinguishes static grounding that presumes understanding from dynamic grounding that builds it?
- How should agents separate planning from perception grounding?
- Why can't static grounding alone close the gap between agreement and understanding?
19 specific questions
- How does externalizing reasoning into harness artifacts improve agent reliability?
- How do agentic systems hide harness failures from benchmarks?
- What components of agent scaffolding most impact domain-specific output quality?
- How do different harness designs produce different agent behaviors from the same model?
- How much realized agent capability comes from the harness versus the model?
- Why does the harness layer accumulate distributed behaviors over time?
- Why does externalized state beat parameter scaling for agent reliability?
50 specific questions
- Why do agents ignore condensed experience in favor of raw data?
- Can agents improve if we constrain how much history they retain?
- Should agents update memory after every turn or batch process sessions?
- How should agents compress episodic interactions into working memory without accumulation?
- Do agents prefer raw experience over condensed summaries of past actions?
- What drives the choice between storing raw episodes versus abstracted rules?
- Does reducing interaction history cost agents performance on their tasks?