Deloitte Billed the Australian Government for an AI Report That Was Never Properly Reviewed
What happened
A major consulting firm delivered a report to the Australian government, collected payment on a six-figure contract, and never told the client that a large language model had generated substantial portions of the content. When the errors embedded in that content surfaced, Deloitte Australia returned part of the fee and revised the document to include the disclosure it had initially withheld.
Deloitte used GPT-4o, deployed through Azure OpenAI, to produce a report for the Department of Employment and Workplace Relations. The tool generated false or misleading information, errors commonly grouped under the term hallucination, and that material made it into the version Deloitte submitted. The company did not mention AI use anywhere in the original document. The choice not to disclose was not a technical oversight; Deloitte later added the disclosure to a revised version, which means the information existed and was left out of the first submission.
The company agreed to return the final installment of the AUD 440,000 contract, framed as a voluntary penalty rather than a court-ordered refund. The Department confirmed that the substantive recommendations were unchanged between the original and revised versions, though that determination rested on Deloitte's own assessment rather than an independent review of the AI-generated sections. The settlement left unanswered the question of how much analytical weight the hallucinated content had actually carried.
The partial repayment did not settle the political reaction. Several politicians called for a full refund and argued that the government needs much stronger requirements around how suppliers deploy AI on public-sector work. The incident raised broader questions about whether existing procurement standards give agencies any reliable mechanism to know what actually produced the analysis they are buying, and whether a voluntary disclosure regime is sufficient to carry the weight now placed on it.
The deeper problem is not that a language model produced errors. Models produce errors routinely, and the expectation that they do is exactly why review and disclosure requirements exist. The problem is that a firm could submit AI-generated work to a government client, collect payment, and face no mandatory obligation to say so until after the errors were already embedded in policy-relevant analysis. That gap is precisely what accountability infrastructure is built to close: a provable record of what a system did, who reviewed its output before it left the firm, and whether the client was informed. Without that record, every government contract that touches a language model is one undisclosed hallucination away from the same outcome.
Reported impact
- Affected parties
- Not publicly disclosed
- Harm type
- Not publicly disclosed
- Scale
- Not publicly disclosed
- Financial impact
- Not publicly disclosed
- Regulatory action
- Not publicly disclosed
Classification
Relevant governance controls
Governance control mapping is not available for this record.
- No controls mapped
Not publicly disclosed
Control mapping is analytical. It does not state that any control would have prevented the incident.
Sources and evidence
This record was researched and written by the Index. The event is also catalogued in the following database, which is listed for cross-reference.