AI in Audit: The Evidence Has to Survive Inspection
AI in audit has quietly crossed a line. A recent Caseware study found AI is now core to US audit firms, with the focus shifting from adoption to control. The largest firms have spent heavily building AI into their audit and tax platforms. And regulators have started writing down what they expect: in October, Singapore's Accounting and Corporate Regulatory Authority issued practice guidance on using AI responsibly in audits, built around integrity, professional judgement, accountability, transparency, and data confidentiality. The common thread is a shift in the question inspectors ask. It is no longer "do you use AI?" It is "show us how you checked what the tool told you."
Where AI Sits in the Audit and What Inspectors Will Ask
Each use of AI creates a different kind of evidence, and each invites a different question on inspection.
| Audit Area | Typical AI Use | Likely Inspection Question |
|---|---|---|
| Risk assessment | Scanning filings, news, and prior files for risk factors | How did the team challenge what the tool missed? |
| Journal entry testing | Scoring every entry for unusual patterns | How were thresholds set and flagged items followed up? |
| Full-population analytics | Testing all transactions instead of a sample | Is the data complete and was the tool validated? |
| Document extraction | Reading contracts, invoices, and confirmations | What was the error rate and how was it checked? |
| Workpaper drafting | Writing memos and summaries | Where is the auditor's own judgement in the file? |
| Technical research | Answering accounting and standards questions | Were citations verified against the actual standard? |
From Adoption to Control
The early debate about AI in audit was about permission. Could firms use it, and would regulators accept the output? That debate is largely over. Tools that score every journal entry, extract terms from thousands of contracts, or draft a first version of a planning memo are now ordinary parts of engagements at firms of every size. Agentic tools that plan and execute multi-step procedures are starting to appear as well.
What has replaced the permission question is a quality question. An audit opinion rests on sufficient, appropriate evidence and on the auditor's professional judgement. When a tool produces part of that evidence or drafts part of that judgement, the firm has to show the tool was fit for purpose and that a qualified person evaluated what it produced.
What AI in Audit Changes About Evidence
Full-population testing is the clearest example. Traditional sampling accepted that some items would go untested, and the methodology accounted for that. When AI tests every transaction, the logic changes. The tool will flag exceptions, sometimes thousands of them, and the audit team now owns every flag. An exception identified and not followed up is worse, from an inspection standpoint, than an exception never seen. More coverage means more responsibility, not less.
What Regulators Are Signalling
Guidance is arriving from several directions. Alongside Singapore's new practice guidance, the UK Financial Reporting Council has published material on AI use in audit, and the US PCAOB's staff have shared observations on generative AI in audit work. The wording differs, but the themes are consistent:
- Judgement stays with the auditor. A tool can inform a conclusion but cannot own it.
- Tools need validation. Firms should understand what a tool does, test it, and know its limitations before relying on it.
- Documentation must show the work. The file should record which tool was used, on what data, and how outputs were evaluated.
- Skepticism applies to AI output. A confident summary deserves the same challenge as a confident client explanation.
- Client data must be protected. Confidential information should not be entered into tools without appropriate safeguards.
Cross-Check a Technical Accounting Answer
Ask six AI models the same standards question and see where their readings of the guidance differ.
Try Talkory FreeWhere AI Goes Wrong on an Engagement
The failures that worry inspectors are rarely dramatic. They are small, plausible, and easy to miss under deadline pressure. A research assistant cites a paragraph of an accounting standard that says something slightly different, or does not exist. An extraction tool misreads a renewal date in a lease, and the error flows into a calculation nobody re-performs. An anomaly model is tuned so conservatively that it flags little, and the team reads silence as comfort.
There is a newer risk too. Clients use AI as well, and explanations for unusual transactions increasingly arrive in fluent, well-structured prose that may itself be machine-generated. A polished explanation is not corroborating evidence. Skepticism has to extend to the form an answer arrives in, not just its content.
Seven Controls That Hold Up Under Inspection
- Maintain an approved tool register. List which AI tools may be used, for which procedures, and on what data.
- Validate before reliance. Test each tool on known data, record error rates, and repeat when the tool or model changes.
- Set follow-up rules for exceptions. Define how flagged items are investigated, by whom, and how resolution is documented.
- Verify every citation. Any reference to a standard, regulation, or ruling produced by AI must be checked against the source text.
- Require reviewer sign-off on AI-drafted content. The reviewer should be able to explain the conclusion without the tool.
- Record tool, version, inputs, and outputs. An inspector should be able to see exactly what the tool did on that engagement.
- Protect client data by design. Use tools that keep confidential information inside approved environments, with clear retention rules.
Pros and Cons of AI in Audit
- Pro: broader coverage. Testing whole populations can surface issues sampling would miss.
- Pro: time for judgement. Less manual extraction leaves more hours for the areas that need experienced thinking.
- Pro: consistency. Standard procedures run the same way across engagements and teams.
- Con: more exceptions to resolve. Full coverage creates follow-up obligations that can swamp a team.
- Con: hidden tool error. Extraction and classification mistakes are hard to spot without re-performance.
- Con: skills erosion. Junior staff who only review AI output may not develop the instincts that catch problems.
Real Scenarios Worth Thinking Through
These scenarios are illustrative, showing how AI in audit plays out in practice rather than presented as verified case studies.
A firm runs journal entry analytics across a client's full ledger and the tool flags several hundred unusual entries. The team investigates the top fifty and documents them well, but the file is silent on the rest. On inspection, the question is not about the fifty. It is about the others, and why there was no documented basis for leaving them.
A senior uses an AI research tool to draft a memo on revenue recognition for a complex contract. The memo cites a paragraph from the standard that, on checking, addresses a different situation. The reviewer catches it because the firm requires every citation to be verified, which is the control working exactly as intended.
An engagement team uses AI to extract terms from hundreds of leases. A sample re-performance finds a small but consistent error in how one lease format is read. Because the firm validated the tool and re-performed a sample, the error is corrected before it affects the lease liability.
Client Data That Must Stay Inside the Firm?
Enterprise plans cover private deployment, custom data residency, dedicated infrastructure, and an SLA.
Talk to Enterprise Sales“After testing multiple AI models on coding, research, and business prompts, combined outputs produced more reliable results than any single model.” Internal multi-model evaluation, Talkory research team.
Documentation Is the Product
Inspectors do not watch an audit happen. They read the file. That makes documentation the place where AI use succeeds or fails from a regulatory point of view. A strong file shows which tool was used and why it was appropriate, what data it ran on and how completeness was confirmed, what it produced, what the team did with each output, and who reviewed the conclusion. A weak file shows a tool output and a tick.
The principles overlap with the record keeping we described in AI audit trails, though the stakes for auditors are sharper because their name is on a public opinion. Tax practitioners face the same dynamic, covered in AI tax preparation: the professional signs, so the professional owns the error.
Why Talkory Wins
Talkory is not an audit tool and does not produce audit evidence. Where it helps is the research and reasoning around an engagement. It runs the same question across GPT, Claude, Gemini, Grok, Perplexity Sonar, and Kimi K3 in one pass, so a team researching an accounting treatment, a disclosure requirement, or the meaning of a standard can see whether six independent models agree. Consensus suggests a mainstream reading worth confirming in the source. Disagreement on a citation, a threshold, or an interpretation is exactly the point where a reviewer should open the standard and consult the technical team. It sharpens the checking that regulators now expect.
Final Verdict
AI in audit is no longer an experiment, and regulators have stopped treating it as one. Their expectations are converging on a simple principle: firms may use AI widely, but the auditor must still exercise judgement, validate the tools, follow up what they flag, verify what they cite, protect client data, and document all of it. The firms that do well on inspection will not be the ones using the most AI. They will be the ones whose files show, clearly and specifically, how every AI output was checked.
Verify Your Next Technical Question
Compare six AI models on the same accounting or standards question before it reaches the file.
Try Talkory FreeFrequently Asked Questions
Can auditors rely on AI-generated evidence?
They can use it when the tool has been validated, the data is complete, and the team has evaluated the output with professional skepticism. The auditor remains responsible for the evidence and the conclusion, whatever tool produced part of it.
What does ACRA's guidance on AI in audit cover?
Singapore's practice guidance on responsible AI use in audit highlights integrity, professional judgement, accountability, transparency, and data confidentiality and security. It reflects themes shared by other audit regulators.
How should audit files document AI use?
Record which tool and version was used, on what data, how completeness was confirmed, what the tool produced, how each relevant output was followed up, and who reviewed the conclusion.
Does full-population testing with AI reduce audit risk?
It can, but it also creates an obligation to follow up the exceptions it flags. Flagged items that are not investigated or documented can create more inspection risk than a well-designed sample.
Can auditors put client data into public AI tools?
Generally they should not without safeguards. Regulators stress confidentiality and data security, so firms should use approved tools that keep client information within controlled environments with clear retention rules.
Get 5 AI perspectives on this topic
Talkory runs your question through GPT, Claude, Gemini, Grok, Sonar & Kimi K3 simultaneously, then cross-checks the answers.