AI in Consulting: The Review Gap Behind the Refunds
AI in consulting has moved from experiment to everyday habit on almost every engagement. Analysts use it to scan literature, summarise interviews, draft sections, and assemble reference lists at a pace no junior team could match. The drafting is not the problem. The problem is what happens when a confident, well-formatted paragraph carries a citation that does not exist, a statistic nobody can trace, or a quote attributed to a judgment that never contained it, and that paragraph travels through three layers of review before landing in front of a paying client.
Traditional vs AI-Assisted Deliverable Risk
The kind of error changes when the first draft comes from a model, and so does the kind of review needed to catch it.
| Factor | Traditional Engagement | AI-Assisted Engagement |
|---|---|---|
| Where a claim comes from | An analyst reads the source directly | A model summarises or recalls the source |
| Typical error | Misread or outdated figure | Invented reference, merged sources, fabricated quote |
| What the error looks like | Often visibly inconsistent | Polished, consistent, entirely plausible |
| Who usually catches it | A manager checking working notes | Often nobody until the client or an outsider reads closely |
| Review time | Roughly proportional to drafting time | Drafting time collapses, review time rarely grows to match |
| Client exposure | Correction and apology | Refund, public withdrawal, reputational damage |
What the Public Refund Cases Actually Show
In 2025 Deloitte Australia agreed to repay part of its fee on a report prepared for a federal government department, on a contract reported at around AU$440,000, after an academic reviewing the document found references that did not exist and a quotation attributed to a court judgment that the judgment did not contain. The corrected version acknowledged that a generative AI tool had been used in preparing it. Further corrections and withdrawals involving other large advisory firms have been reported since, and the pattern across them is remarkably consistent.
None of these reports failed because the analysis was careless or the conclusions were absurd. They failed at the evidence layer. The recommendations were broadly defensible, the writing was professional, and the formatting met house standards. What broke was the thread between a sentence and the thing it claimed to rest on. That thread is exactly what language models are worst at preserving, and exactly what a time-pressed senior reviewer is least likely to test.
Why AI in Consulting Errors Survive Partner Review
Consulting quality control was designed around a particular failure profile. Junior staff make errors of judgement, omission, and arithmetic, and experienced reviewers are very good at spotting those because they have seen them hundreds of times. A partner reading a draft asks whether the argument holds, whether the recommendation fits the client, and whether anything feels off. Realistically, they are not opening every footnote.
The Fluency Problem
Fabricated references rarely look fabricated. They carry plausible author names, credible journal titles, realistic years, and page ranges. An invented quote from a court decision reads like legal language because the model has absorbed an enormous amount of legal language. The signal reviewers depend on, something that looks slightly wrong, is precisely the signal these tools remove.
Why AI in Consulting Errors Look Like Diligence
There is an uncomfortable irony here. A report with sixty references looks more thorough than one with fifteen. AI in consulting makes long reference lists cheap to produce, and a long reference list reads as diligence to reviewers and clients alike. The volume of apparent evidence rises while the proportion of verified evidence can quietly fall, and nothing in a normal review process measures that ratio. Law firms learned the same lesson in public, where lawyers were fined for fake citations that had passed their own internal checks.
Six Places Fabrication Enters a Client Report
These entry points recur across strategy, public sector, and operations work.
- Literature and market scans. Asking a model to find supporting sources is the single most common route to references that do not exist.
- Reference list clean-up. Tools that tidy citation formatting can silently change authors, years, or titles, turning a real source into a phantom one.
- Quotations. Paraphrases get promoted into direct quotes, and direct quotes get attributed to the wrong document or the wrong speaker.
- Statistics without a trail. A figure copied from a draft summary into a slide loses its origin within two iterations of the deck.
- Legal and regulatory summaries. Descriptions of what a law, ruling, or standard requires are fluent and frequently imprecise.
- Benchmarks and examples. Illustrative company examples drift into case studies presented as fact, complete with invented detail.
Check Every Claim Against Six Models
Paste a paragraph and see which statements the models agree on before it goes to the client.
Try Talkory FreePros and Cons of AI in Client Work
None of this is an argument against the tools. It is an argument about where the effort has to move.
- Pro: faster first drafts. Teams reach a reviewable draft in a fraction of the time, which frees hours for thinking rather than typing.
- Pro: broader early scans. A model can surface angles and sources that a small team would never have time to find.
- Pro: better synthesis of interviews. Summarising dozens of stakeholder conversations into themes is a genuine strength.
- Con: evidence becomes untraceable. Claims detach from their sources unless the workflow forces them to stay attached.
- Con: verification does not get faster. Drafting sped up, but confirming a reference still takes as long as it always did.
- Con: liability stays with the firm. The engagement letter names the firm, not the software, and clients have shown they will ask for money back.
Real Scenarios Worth Thinking Through
These scenarios are illustrative, showing how AI in consulting plays out in practice rather than presented as verified case studies.
A strategy team prepares a market entry report and includes a market size figure attributed to an industry association. The association never published that number. The client repeats it in a board presentation, and a competitor who knows the association's actual research challenges it in a public forum.
A public sector review cites three academic papers behind a key finding. Two are real. The third combines genuine authors and a genuine journal with a title nobody ever wrote. An academic reading the published report finds it within a day, and the story becomes the fake paper rather than the finding.
An operations diagnostic uses AI to summarise forty interviews. The summaries are broadly accurate, but one memorable quote in the final deck merges two different interviewees. The client recognises one of them and asks, in the steering meeting, who actually said it.
Keep Client Data Inside Your Own Boundary
Enterprise plans cover private deployment, custom data residency, dedicated infrastructure, and an SLA.
Talk to Enterprise Sales“After testing multiple AI models on coding, research, and business prompts, combined outputs produced more reliable results than any single model.” Internal multi-model evaluation, Talkory research team.
A Deliverable Review Standard That Holds Up
Banning the tools is neither realistic nor necessary. What works is separating two jobs that used to be merged. Quality review asks whether the argument is sound. Source verification asks whether every factual claim rests on something real. On an AI-assisted engagement those need different people, or at least different passes, because the second job no longer happens by accident while doing the first.
Three rules carry most of the weight. Every factual claim in a client deliverable must have a named source that someone can actually open. One person who did not write the section is accountable for verifying those sources, and their name is recorded. Sections drafted with AI are flagged internally so reviewers know where to look hardest. Then check by risk: every direct quote, every legal or regulatory statement, and every statistic in the executive summary gets verified in full, while supporting references in the body are sampled. It sounds bureaucratic until the first client asks. The mechanics of checking references quickly are covered in our consensus method for citation accuracy.
Should Consultants Tell Clients They Used AI?
Increasingly, yes, and many client procurement teams now ask directly in their questionnaires. Disclosure is not an admission of weakness. What a client cares about is that the evidence is real and that someone accountable checked it. A short statement describing how AI was used and how its output was verified tends to build more confidence than silence, and silence becomes very expensive if a problem surfaces later.
The pricing conversation follows close behind. If AI shortens drafting, clients will reasonably expect that to show up somewhere. Firms that reinvest saved hours into verification have a far better answer to that question than firms that simply banked the time.
Why Talkory Wins
Ask a single model whether a source exists and it will often say yes, confidently, particularly if it produced the source in the first place. Talkory sends the same claim to GPT, Claude, Gemini, Grok, Perplexity Sonar, and Kimi K3 together. A real, established source tends to be recognised consistently across independently trained models. A fabricated one tends to produce disagreement, hedging, or different details from each model, and that disagreement is a fast triage signal for a reviewer.
It does not replace opening the source. It tells you which of the sixty references to open first, which turns an impossible review into a manageable one on a real deadline.
Final Verdict
AI in consulting is not going away, and it should not. It makes teams faster and often sharper. What the public refund cases show is that the risk sits in one narrow place: the connection between a claim and its evidence. Firms that assign a named verifier, check every quote and every headline statistic, triage references with cross-model disagreement, and tell clients how the work was checked will keep the speed without the refunds. Firms that treat a polished draft as a verified one are borrowing against their reputation.
Frequently Asked Questions
Can consulting firms be forced to refund fees over AI errors?
It depends on the contract, but it has already happened. A large firm agreed to repay part of a government contract after its report was found to contain fabricated references linked to AI use. Engagements typically promise professional care, and invented evidence is very hard to defend under that standard.
How do fabricated references get past internal review?
Review processes were built to catch errors of judgement rather than invented evidence. Fabricated citations look plausible, with real-sounding authors, journals, and page numbers, and senior reviewers rarely open every footnote. The polish that makes AI drafts attractive also removes the visual cues reviewers normally rely on.
Which parts of a consulting report need the most checking?
Direct quotations, legal and regulatory statements, statistics in the executive summary, and any reference supporting a key finding. These carry the most client and reputational risk, and they are also where AI tools most often invent or distort detail.
Should consultants disclose AI use to clients?
Disclosure is becoming expected, and many client procurement teams now ask about it directly. Explaining how AI was used and how its outputs were verified tends to build confidence, while undisclosed use becomes a serious trust problem if an error surfaces later.
Does using several AI models help verify a report?
It helps with triage. Real, established sources are usually recognised consistently across independent models, while fabricated ones tend to trigger disagreement or inconsistent details. That points reviewers to the references most worth opening, although the source itself must still be checked.
Get 5 AI perspectives on this topic
Talkory runs your question through GPT, Claude, Gemini, Grok, Sonar & Kimi K3 simultaneously, then cross-checks the answers.