Can AI quality control catch fabricated citations in professional reports?
This explores whether existing review processes at major firms can catch AI-generated errors like fake quotes and nonexistent citations before client delivery. It matters because firms claim to maintain quality oversight while adopting AI tools.
CFO Dive reports that Deloitte Australia refunded "over 97,000 Australian dollars ($63,000 USD)" to Australia's Department of Employment and Workplace Relations (DEWR) after a paid assurance report it delivered was found to contain "artificial intelligence-generated errors." The refund covers less than a quarter of the roughly A$440,000 contract, executed in December 2024, for a review of the IT system the department uses to automate welfare penalties. According to the Associated Press, as cited by CFO Dive, the original report contained "a fabricated quote from a federal court judgment and references to nonexistent academic research papers."
DEWR's spokesperson said Deloitte "conducted the independent assurance review and has confirmed some footnotes and references were incorrect," while adding that "the substance" of the review had been retained — drawing a line between the report's citations and footnotes, where the AI-generated failure occurred, and its substantive findings, which the client maintains survived. Hofstra accounting professor Jack Castonguay frames the failure as a training and quality-control gap rather than a one-off glitch, telling CFO Dive that firms "need to not only train staff on how to use AI effectively, but they must also train them on ethical use and maintaining quality control," and that AI output "must be reviewed as if it were prepared by an intern or new hire."
The incident is a concrete instance of the failure mode described in Can organizations lose scrutiny capacity while keeping oversight forms?: Deloitte presumably had some review process in place for a government-facing assurance report, yet fabricated citations reached the delivered, paid document, meaning the scrutiny step that should have caught invented quotes and papers did not function as intended. It also sits in tension with How close are frontier AI models to expert work quality?: that benchmark finds frontier models approaching expert quality on judged work tasks, but here a live, paid professional deliverable from a major firm shipped with fabricated legal and academic citations, pointing to a gap between benchmarked capability and quality control under real client conditions. The public exposure and partial refund also give the Does disclosing AI use damage how trustworthy you seem? dynamic a firm-level, involuntary analog: credibility costs that Schilke and Reimann find at the level of an individual choosing to disclose AI use appear here as a reputational and contractual cost imposed on a firm after AI-generated errors were discovered by others.
The excerpt does not establish how the errors were produced — which AI tool was used, whether its use was disclosed to DEWR in advance, or whether the review failure traces to time pressure, understaffing, or a specific breakdown in Deloitte's checking process. It also gives no basis for estimating how common this kind of failure is across Deloitte's or other Big Four firms' AI-assisted deliverables, since this is one disclosed case that became public rather than a sampled or audited rate. The narrow, supportable implication is that fabricated AI output can reach a paid, government-facing professional deliverable despite an "independent assurance review," and that when caught, it carries a direct financial consequence; whether review processes industry-wide are adequate to catch such errors before delivery is not something this source speaks to.
Inquiring lines that read this note 5
This note is a source for these research framings, grouped by the broader line of inquiry each explores. Scan the bold lines of inquiry; follow any specific question forward.
What are the real-world consequences of AI citation hallucinations? How do hallucinated citations emerge in AI scholarly output?Related concepts in this collection 4
This note in its neighbourhood — explore the map, then jump to a related concept in the list below.
Click a node to walk · click center to open · click Open in graph to see this note in the full knowledge graph
-
Can organizations lose scrutiny capacity while keeping oversight forms?
When human review steps remain in organizational processes, do they retain meaningful scrutiny ability or can that capacity erode invisibly? This matters because paper oversight looks identical to real oversight in audits.
this incident is a real case of the review step failing to catch fabricated AI content before delivery
-
How close are frontier AI models to expert work quality?
GDPval benchmarked frontier models on 1,320 expert-built tasks across 44 occupations, using head-to-head expert judgment to measure whether AI is approaching human deliverable quality in knowledge work.
contrasts benchmarked near-expert quality with a real paid deliverable that still contained fabricated citations
-
Does disclosing AI use damage how trustworthy you seem?
When people learn you used AI to create work, do they trust you less? Schilke and Reimann tested this across 13 experiments with over 5,000 participants to understand whether transparency about AI reliance backfires.
the public exposure and refund show a firm-level credibility cost analogous to the individual trust penalty Schilke and Reimann measure
-
Can AI-assisted reports pass quality checks with fabricated citations?
A Deloitte government report contained over a dozen invented references and fake quotes that escaped internal review. The question is whether AI-generated content can systematically bypass citation verification in high-stakes professional work.
Extends: identifies the academic who found the fabrications and the Azure GPT-4o tool chain Deloitte used to produce the report
Related papers in this collection 8
Papers most semantically related to this note, ranked by cosine similarity in the embedding space.
- Deloitte to refund government after using AI in $440k report
- Your AI Strategy Advisor Is Giving Everyone the Same Advice
- Deloitte refunds Australian government for AI-flawed report
- Pangram Predicts 21% of ICLR Reviews are AI-Generated
- Machine Bullshit: Characterizing the Emergent Disregard for Truth in Large Language Models
- Stranded Credentials: Keeping Online Reputation Systems Informative in the AI Era
- Tortured phrases: A dubious writing style emerging in science. Evidence of critical issues affecting established journals
- People Overtrust AI-Generated Medical Advice despite Low Accuracy
Original note title
Deloitte refunds the Australian government after its AI-assisted report fabricated a court quote and nonexistent research papers