Does AI theme-mapping perform as well as human reviewers?
Can an AI system match expert judgment when automatically sorting consultation responses into themes? Understanding this matters for scaling policy feedback analysis without sacrificing accuracy.
UK Incubator for AI (i.AI), the government unit that built Consult, reports that in a live Scottish Government consultation the tool's AI-generated theme mappings agreed with expert reviewers almost as closely as the reviewers agreed with each other. Consult scored a mean multi-label F1 of 0.76 (95% CI 0.75–0.77) against reviewer-assigned themes, and reviewers left 60% of responses' Consult-assigned themes unchanged. i.AI compares this to a prior evaluation where two human reviewers agreed with each other at F1 0.81 (95% CI 0.78–0.83), and in this evaluation reviewers matched each other's labels only 62% of the time — "roughly the same as Consult's performance." The report concludes "differences between Consult and the expert reviewers had minimal effects on the overall theme rankings," which it treats as more policy-relevant than exact counts, since identifying the top themes "is usually more important than knowing the exact number of times the theme appeared."
Consult runs two stages with a "human in the loop" at each: "Theme Generation," where topic modelling over the full response set proposes themes that expert reviewers sign off on, and "Theme Mapping," where an LLM assigns themes to each response, which reviewers then verify or amend (median 23 seconds per response). i.AI frames the comparison as inherently fuzzy rather than a clean accuracy score, "given theme mapping is inherently subjective" and because "there is no objective truth for what a correct theme is" — the benchmark is one reviewer's judgment, not ground truth. Where reviewers diverged from Consult, they added themes more than twice as often as they removed them (1,671 vs. 763), and used the open-ended "Other" label nearly 17 times more often than Consult did (about 800 vs. 47) — evidence that the generated theme list, not the mapping step, was the weaker point: "Consult struggled to identify missing themes."
This sits alongside Can clinicians tell GPT-4 advice apart from expert advice? as another case of near-parity between an LLM-based system and expert human judgment on a structured task, but the mechanism i.AI gives is different: it isn't that Consult is mistaken for an expert, it's that mapping open text to themes carries enough inherent subjectivity that two humans disagree about as much as a reviewer and Consult do. It parallels What tasks do expert data storytellers trust to LLMs? more directly: Consult's design puts generation and mapping with the model while judgment and sign-off stay with reviewers, and i.AI's own account of reviewer friction — sign-off described as "cognitively demanding and time intensive," reviewers wanting more confidence before reviewing only a share of responses — shows the tool bought verification labor, not elimination of review.
The evaluation covers one live consultation on one topic, benchmarks Consult against a single reviewer's labels rather than a true consensus, and is run and reported by the team that built Consult and states its mission as expanding AI use across government ("harness the opportunity of AI for public good") — this is i.AI evaluating its own product toward wider adoption, not an independent audit. It does not establish that an F1 near 0.76 generalizes to consultations with different response styles, volumes, or topics, nor that the 60%-unchanged rate would hold without the four-to-six-reviewer team used here. What it supports, at the strength given, is narrower: for this consultation, Consult's discrepancies with reviewers were concentrated in a few themes and in the generation stage rather than spread evenly across mapping, and did not move the policy-relevant theme rankings.
Inquiring lines that read this note 10
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
Do AI coding tools measurably improve developer productivity and code quality?- How do non-experts evaluate AI-generated outputs when they lack implementation expertise?
- Why haven't AI agents replaced human code review workflows?
Related concepts in this collection 3
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can clinicians tell GPT-4 advice apart from expert advice?
This study explores whether trained clinicians can distinguish AI-generated psychological advice from expert advice, and how they rate the quality and empathy of each. The question matters for understanding whether AI might reliably supplement human expertise in mental health settings.
similar near-parity between an AI system and expert judgment, but subjectivity, not mistaken identity, explains the gap here
-
What tasks do expert data storytellers trust to LLMs?
Expert visual data storytellers make strategic choices about which narrative work to delegate to LLMs and which to protect. Understanding these boundaries reveals how human judgment and automation can coexist in knowledge work.
same division of labor: AI handles generation and mapping, reviewers keep sign-off and verification, which stayed effortful
-
Can AI systems safely replace human peer reviewers?
Explores whether AI reviewers meet two critical conditions for automation: maintaining diverse perspectives and resisting score manipulation. Tests whether current systems are ready to handle peer review at scale.
Contradicts: this paper finds AI reviewers show a hivemind effect, agreeing with each other more than humans do — unlike Consult's human-level divergence
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Consult Evaluation: Scottish Government's Non-surgical Cosmetic Procedures Consultation
- AI-Assisted Peer Review at Scale: The AAAI-26 AI Review Pilot
- Stop Automating Peer Review Without Rigorous Evaluation
- Can LLM feedback enhance review quality? A randomized study of 20K reviews at ICLR 2025
- Artificial intelligence vs. human expert: Licensed mental health clinicians' blinded evaluation of AI-generated and expert psychological advice
- Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers
- AI Research Agents Narrow Scientific Exploration
- The Ideation-Execution Gap: Execution Outcomes of LLM-Generated versus Human Research Ideas
Original note title
i.AI finds Consult's theme mapping diverged from reviewers about as much as reviewers diverged from each other — and rarely changed the top themes