Available but Unclaimed: An Empirical Study of Human-AI Synergy

Paper · arXiv 2609.16793 · Published September 15, 2026
Reasoning Architectures

People increasingly reason with large language models (LLMs), yet complementary capabilities do not guarantee outperforming both components. In a between-subjects study, participants (N=535) solved a 40-item battery of matrix reasoning, mental rotation, syllogisms, and letter-string analogies, unaided or with GPT-5.6-Luna, Claude Opus 4.8, Gemini 3.6 Flash, or Kimi K3. Each assisted trial required consultation with the model. Each model answered every item alone 100 times under matched elicitation. The assisted–unaided accuracy difference increased with item-level LLM competence. Deference varied across tasks and increased with competence within tasks. Post-advice confidence distinguished correct from incorrect answers less strongly than unaided confidence. In a reference comparison, about half the increase in LLM accuracy carried through to assisted accuracy. How much of that accuracy gain reached participants differed across the models. These findings motivate evaluating LLMs in interaction with humans and designing support for selective deference that preserves independent reasoning.

Introduction. Large language models (LLMs) are unevenly reliable within a single domain. A model may solve one reasoning problem on nearly every attempt and a problem of the same form at chance [42, 49], while its replies need not reveal the difference [83]. Whether a person can distinguish correct from incorrect advice and rely on the model accordingly matters for effective human–AI interaction (HAI). People and LLMs exhibit different strengths and weaknesses when working on a problem. One can theoretically compensate for the other’s errors [27, 72]. Complementarity denotes this potential, created where human and AI errors differ. Synergy denotes the realized case in which the team actually outperforms both components [73]. During an interaction, answers can change as the conversation proceeds. Differing errors therefore create an opportunity, but do not guarantee that the person can recognize and correct them. We call the realized share of independent-error reference headroom synergy capture. However, achieving synergy is not built into an LLM’s configuration.

Discussion / Conclusion. We asked when a person and an assistant working together outperform both components. Assisted accuracy fell below the item-wise better-component reference, while the battery-level assisted–LLM difference remained uncertain. An AI assistant is not uniformly reliable. On two items of the same kind, it can be near-certain on one and no better than guessing on the next. We measured that variation by running each assistant repeatedly on the same items. Participants’ deference varied substantially by task and also increased with item competence. Consistent with Vaccaro et al. [73], we found no clear advantage over the better component. Riedl and Weidmann’s modeled AI benefit instead compares assisted with unaided performance [63]. For AI development. These data compare solo and assisted performance on the same items. Joint evaluation has both conceptual [26] and empirical precedents [14, 63], while benchmark construct validity remains a concern [4]. Pass-through adds a measure of how assisted accuracy varies with item competence under a specified protocol.

Lines of inquiry this paper opens 24

Research framings built by reading the notes related to this paper — the questions it feeds into.

Does AI assistance promote real skill development or substitute for independent learning? How do capability benchmark scores systematically misrepresent true model abilities? How can infrastructure records verify actual agent behavior? What training dynamics and scale trigger emergence of reasoning capabilities? What fundamental constraints limit how effectively agents can improve themselves? How should designers communicate what AI systems truly are and can do? What linguistic features distinguish AI-generated text from human writing most reliably? How does evaluation scope and dimensionality affect what we measure? How does the generation-verification gap limit what we can measure about AI reasoning? What causes reasoning models to fail or wander off track? Can models improve accuracy without degrading reasoning quality? What explains language models' asymmetric difficulty with implicit versus explicit linguistic relations? Is reasoning capability latent in base models or created by post-training? Do language models reason like humans or mimic surface patterns? Do language models learn genuine understanding or just surface patterns? Why don't LLMs reliably translate capability into accurate outputs? Should AI communication design follow human conversation norms or develop distinct machine-specific principles?