A top consulting firm's AI-assisted report had fake quotes and fake sources — and still passed review before anyone caught it.
How common are undetected AI fabrication errors across Big Four consulting firms?
This explores how often AI-generated fabrications like fake citations, invented quotes and made-up sources slip past review at large professional-services firms such as Deloitte, PwC, EY and KPMG. Because the corpus can't give a frequency, it asks what the documented cases reveal about why these errors go unnoticed.
This explores how widespread undetected AI fabrication is inside Big Four consulting work. The direct answer is that the corpus has no prevalence data. It documents one public case, at Deloitte, and nothing that compares firms or counts incidents. In that case Deloitte refunded A$97,000 to the Australian government after delivering an assurance report containing fabricated court quotes and fake academic references Can AI quality control catch fabricated citations in professional reports?. The report went through review before it shipped, and the fabrications got through anyway. One incident can't tell you how common this is. It does show that a firm's nominal human oversight is not a reliable filter.
The more useful question may be why we can't know how common it is, and here the corpus has a lot to say. Fabrications in professional reports are hard to count because they are designed, in effect, to look right. A citation that is invented but plausible doesn't stand out. It gets through because it matches what a reviewer expects to see. Research on deployed AI systems argues that the failures that matter most are "plausible rather than shocking, distributed rather than localized, normalized by workflows". They stay hidden less because they are technically obscure and more because our checking habits assume errors will look like errors Why do safety failures remain invisible to our evaluation methods?. A related argument holds that more automation produces more polished output, and polish hides mistakes rather than removing them. On that view, integrity becomes a question of disclosure and accountability rather than better detection tools Does more automation actually hide rather than eliminate errors?.
Software development offers a useful comparison. Stack Overflow's 2025 survey found that 80% of developers use AI tools, while trust in their accuracy fell to 29%. The top complaint was code that looks correct but contains subtle errors Why do developers keep using AI tools they don't trust?. Consulting deliverables likely face the same split between how much people use AI and how much they trust it, with one important difference. Code has tests and compilers that eventually expose the problem. A footnote in a policy report usually has nothing like that. If AI does best at tasks whose answers are easy to check Does task verifiability determine what AI systems will learn to solve?, then reference-heavy advisory writing is close to the worst case: checking it is slow, done by hand, and often skipped.
Human psychology adds to the problem. Fluent AI text encourages people to treat a convincing-sounding claim as if it were verified fact. That trap gets worse when the reviewer already agrees with the report's conclusions Why do people trust AI outputs they shouldn't?. A reviewer at a deadline, skimming a well-written draft that supports the client's expected findings, is in exactly the situation this research describes.
So the honest answer to "how common?" is that nobody is measuring it in a way the corpus captures. One paper notes that we lack any instrument for whether AI errors stay visible and recoverable across a whole organization. There are only partial measures of separate pieces How can we measure whether AI errors stay visible and recoverable?. The Deloitte case came to light because an outside reader checked the citations. That suggests the cases we know about depend on someone outside the firm checking, not on how often fabrications actually occur.
Sources 7 notes
Deloitte refunded A$97,000 after delivering a government assurance report containing fabricated court quotes and fake academic references. The incident reveals that nominal human oversight did not catch AI-generated errors before a paid deliverable reached the client.
Deployed AI systems fail in ways that our instruments cannot see: plausible rather than shocking, distributed rather than localized, normalized by workflows rather than immediately legible. The problem is not mystery but mismatched assumptions about failure shape.
Greater automation produces polished outputs that hide errors rather than eliminate them. Scientific integrity therefore depends on disclosure, accountability, and human-governed collaboration—not better fabrication detection tools.
Stack Overflow's 2025 survey shows 80% of developers use AI tools while trust in accuracy fell from 40% to 29%. The primary complaint: AI code that looks correct but contains subtle errors, creating a verification burden that erodes confidence faster than usage grows.
Wei argues that AI solves tasks proportional to how easily solutions can be verified, and that verifiability gaps can be narrowed by pre-investing in answer keys, test suites, or measurement infrastructure. This mechanism explains RL's effectiveness across domains from sudoku to molecular discovery.
Show all 7 sources
Rose-Frame identifies map-territory confusion, intuition-reason conflation, and confirmation-bias reinforcement as traps that multiply their distorting effects when they co-occur. Evidence from cross-linguistic overreliance and architectural transformer biases confirms the compounding mechanism operates universally.
Partial instruments exist for individual conditions in isolated settings, but none measures the full socio-technical system the paper identifies as necessary. Visibility has a model-side measure (chain-of-thought disclosure), containment has incident-level counts, and recoverability has rollback timing, yet none bridges all four or captures human-institution factors.
Papers this line draws on 8
The research behind the notes this line reads — ranked by how closely each paper relates.
- The safety failures we are not instrumenting: a perspective on hidden safety-critical challenges in modern AI systems
- Sycophancy Towards Researchers Drives Performative Misalignment
- On the Reasoning Capacity of AI Models and How to Quantify It
- Beyond Fixed Representations: The Vocabulary and Verifier Gaps in Open-Ended AI
- How AI Can Degrade Human Performance in High-Stakes Settings
- Evaluation Awareness Is Not One Capability: Evidence from Open Language Models
- The Evaluation Differential: When Frontier AI Models Recognise They Are Being Tested
- Evaluation Awareness in Language Models Has Limited Effect on Behaviour