AI in Audit: Regulators Now Ask How You Checked It

AI in audit is now core practice. What regulators expect when firms use AI on audit evidence, and the controls that let a partner sign with confidence.

AI in Audit: The Evidence Has to Survive Inspection

Quick Answer: AI in audit is now routine, from journal entry testing to drafting workpapers. Regulators are no longer asking whether firms use it. They ask how outputs were checked, how judgement was applied, and whether the file shows it. Firms need tool validation, review standards, and clear records.

AI in audit has quietly crossed a line. A recent Caseware study found AI is now core to US audit firms, with the focus shifting from adoption to control. The largest firms have spent heavily building AI into their audit and tax platforms. And regulators have started writing down what they expect: in October, Singapore's Accounting and Corporate Regulatory Authority issued practice guidance on using AI responsibly in audits, built around integrity, professional judgement, accountability, transparency, and data confidentiality. The common thread is a shift in the question inspectors ask. It is no longer "do you use AI?" It is "show us how you checked what the tool told you."

Where AI Sits in the Audit and What Inspectors Will Ask

Each use of AI creates a different kind of evidence, and each invites a different question on inspection.

Audit AreaTypical AI UseLikely Inspection Question
Risk assessmentScanning filings, news, and prior files for risk factorsHow did the team challenge what the tool missed?
Journal entry testingScoring every entry for unusual patternsHow were thresholds set and flagged items followed up?
Full-population analyticsTesting all transactions instead of a sampleIs the data complete and was the tool validated?
Document extractionReading contracts, invoices, and confirmationsWhat was the error rate and how was it checked?
Workpaper draftingWriting memos and summariesWhere is the auditor's own judgement in the file?
Technical researchAnswering accounting and standards questionsWere citations verified against the actual standard?

From Adoption to Control

The early debate about AI in audit was about permission. Could firms use it, and would regulators accept the output? That debate is largely over. Tools that score every journal entry, extract terms from thousands of contracts, or draft a first version of a planning memo are now ordinary parts of engagements at firms of every size. Agentic tools that plan and execute multi-step procedures are starting to appear as well.

What has replaced the permission question is a quality question. An audit opinion rests on sufficient, appropriate evidence and on the auditor's professional judgement. When a tool produces part of that evidence or drafts part of that judgement, the firm has to show the tool was fit for purpose and that a qualified person evaluated what it produced.

What AI in Audit Changes About Evidence

Full-population testing is the clearest example. Traditional sampling accepted that some items would go untested, and the methodology accounted for that. When AI tests every transaction, the logic changes. The tool will flag exceptions, sometimes thousands of them, and the audit team now owns every flag. An exception identified and not followed up is worse, from an inspection standpoint, than an exception never seen. More coverage means more responsibility, not less.

What Regulators Are Signalling

Guidance is arriving from several directions. Alongside Singapore's new practice guidance, the UK Financial Reporting Council has published material on AI use in audit, and the US PCAOB's staff have shared observations on generative AI in audit work. The wording differs, but the themes are consistent:

  • Judgement stays with the auditor. A tool can inform a conclusion but cannot own it.
  • Tools need validation. Firms should understand what a tool does, test it, and know its limitations before relying on it.
  • Documentation must show the work. The file should record which tool was used, on what data, and how outputs were evaluated.
  • Skepticism applies to AI output. A confident summary deserves the same challenge as a confident client explanation.
  • Client data must be protected. Confidential information should not be entered into tools without appropriate safeguards.

Cross-Check a Technical Accounting Answer

Ask six AI models the same standards question and see where their readings of the guidance differ.

Try Talkory Free

Where AI Goes Wrong on an Engagement

The failures that worry inspectors are rarely dramatic. They are small, plausible, and easy to miss under deadline pressure. A research assistant cites a paragraph of an accounting standard that says something slightly different, or does not exist. An extraction tool misreads a renewal date in a lease, and the error flows into a calculation nobody re-performs. An anomaly model is tuned so conservatively that it flags little, and the team reads silence as comfort.

There is a newer risk too. Clients use AI as well, and explanations for unusual transactions increasingly arrive in fluent, well-structured prose that may itself be machine-generated. A polished explanation is not corroborating evidence. Skepticism has to extend to the form an answer arrives in, not just its content.

Seven Controls That Hold Up Under Inspection

  1. Maintain an approved tool register. List which AI tools may be used, for which procedures, and on what data.
  2. Validate before reliance. Test each tool on known data, record error rates, and repeat when the tool or model changes.
  3. Set follow-up rules for exceptions. Define how flagged items are investigated, by whom, and how resolution is documented.
  4. Verify every citation. Any reference to a standard, regulation, or ruling produced by AI must be checked against the source text.
  5. Require reviewer sign-off on AI-drafted content. The reviewer should be able to explain the conclusion without the tool.
  6. Record tool, version, inputs, and outputs. An inspector should be able to see exactly what the tool did on that engagement.
  7. Protect client data by design. Use tools that keep confidential information inside approved environments, with clear retention rules.

Pros and Cons of AI in Audit

  • Pro: broader coverage. Testing whole populations can surface issues sampling would miss.
  • Pro: time for judgement. Less manual extraction leaves more hours for the areas that need experienced thinking.
  • Pro: consistency. Standard procedures run the same way across engagements and teams.
  • Con: more exceptions to resolve. Full coverage creates follow-up obligations that can swamp a team.
  • Con: hidden tool error. Extraction and classification mistakes are hard to spot without re-performance.
  • Con: skills erosion. Junior staff who only review AI output may not develop the instincts that catch problems.

Real Scenarios Worth Thinking Through

These scenarios are illustrative, showing how AI in audit plays out in practice rather than presented as verified case studies.

A firm runs journal entry analytics across a client's full ledger and the tool flags several hundred unusual entries. The team investigates the top fifty and documents them well, but the file is silent on the rest. On inspection, the question is not about the fifty. It is about the others, and why there was no documented basis for leaving them.

A senior uses an AI research tool to draft a memo on revenue recognition for a complex contract. The memo cites a paragraph from the standard that, on checking, addresses a different situation. The reviewer catches it because the firm requires every citation to be verified, which is the control working exactly as intended.

An engagement team uses AI to extract terms from hundreds of leases. A sample re-performance finds a small but consistent error in how one lease format is read. Because the firm validated the tool and re-performed a sample, the error is corrected before it affects the lease liability.

Client Data That Must Stay Inside the Firm?

Enterprise plans cover private deployment, custom data residency, dedicated infrastructure, and an SLA.

Talk to Enterprise Sales
“After testing multiple AI models on coding, research, and business prompts, combined outputs produced more reliable results than any single model.” Internal multi-model evaluation, Talkory research team.

Documentation Is the Product

Inspectors do not watch an audit happen. They read the file. That makes documentation the place where AI use succeeds or fails from a regulatory point of view. A strong file shows which tool was used and why it was appropriate, what data it ran on and how completeness was confirmed, what it produced, what the team did with each output, and who reviewed the conclusion. A weak file shows a tool output and a tick.

The principles overlap with the record keeping we described in AI audit trails, though the stakes for auditors are sharper because their name is on a public opinion. Tax practitioners face the same dynamic, covered in AI tax preparation: the professional signs, so the professional owns the error.

Why Talkory Wins

Talkory is not an audit tool and does not produce audit evidence. Where it helps is the research and reasoning around an engagement. It runs the same question across GPT, Claude, Gemini, Grok, Perplexity Sonar, and Kimi K3 in one pass, so a team researching an accounting treatment, a disclosure requirement, or the meaning of a standard can see whether six independent models agree. Consensus suggests a mainstream reading worth confirming in the source. Disagreement on a citation, a threshold, or an interpretation is exactly the point where a reviewer should open the standard and consult the technical team. It sharpens the checking that regulators now expect.

Final Verdict

AI in audit is no longer an experiment, and regulators have stopped treating it as one. Their expectations are converging on a simple principle: firms may use AI widely, but the auditor must still exercise judgement, validate the tools, follow up what they flag, verify what they cite, protect client data, and document all of it. The firms that do well on inspection will not be the ones using the most AI. They will be the ones whose files show, clearly and specifically, how every AI output was checked.

Verify Your Next Technical Question

Compare six AI models on the same accounting or standards question before it reaches the file.

Try Talkory Free

Frequently Asked Questions

Can auditors rely on AI-generated evidence?

They can use it when the tool has been validated, the data is complete, and the team has evaluated the output with professional skepticism. The auditor remains responsible for the evidence and the conclusion, whatever tool produced part of it.

What does ACRA's guidance on AI in audit cover?

Singapore's practice guidance on responsible AI use in audit highlights integrity, professional judgement, accountability, transparency, and data confidentiality and security. It reflects themes shared by other audit regulators.

How should audit files document AI use?

Record which tool and version was used, on what data, how completeness was confirmed, what the tool produced, how each relevant output was followed up, and who reviewed the conclusion.

Does full-population testing with AI reduce audit risk?

It can, but it also creates an obligation to follow up the exceptions it flags. Flagged items that are not investigated or documented can create more inspection risk than a well-designed sample.

Can auditors put client data into public AI tools?

Generally they should not without safeguards. Regulators stress confidentiality and data security, so firms should use approved tools that keep client information within controlled environments with clear retention rules.

CK

Chetan Kajavadra, Lead AI Researcher, Talkory.ai

Chetan specialises in AI model evaluation, enterprise AI risk, and multi-LLM orchestration strategy. Reviewed by Mital Bhayani, AI Researcher and SaaS Growth Specialist at Talkory.ai. Connect on LinkedIn →

๐Ÿค–

Get 5 AI perspectives on this topic

Talkory runs your question through GPT, Claude, Gemini, Grok, Sonar & Kimi K3 simultaneously, then cross-checks the answers.

Try Talkory.ai free โ†’
โ† Back to all articles

Related Articles

๐ŸขProfessional Services

AI Tax Preparation: The Preparer Still Owns the Error

The IRS did not ban AI in tax practice. It said something more consequential: diligence, competence, and confidentiality are unchanged, and relying on a tool without checking its work does not meet them.

Read article โ†’
๐Ÿ“ฐAI and Media

Can AI Spot Fake News? We Tested All 5 Models

We built a 20-headline test, half real and half fake, and ran it through ChatGPT, Claude, Gemini, Grok, and Perplexity. Claude scored 90%. Grok scored 70% while sounding 95% confident. Confidence without accuracy is the failure mode that actually spreads misinformation.

Read article โ†’
โœˆ๏ธAI Travel

Best AI for Travel Planning: We Tested All 5 Models

We gave all five AI models the same Tokyo prompt and audited every restaurant, museum, and transit direction. Perplexity scored 95%. Grok scored 63%. A hallucinated restaurant ruins a vacation. Here is what the field looks like.

Read article โ†’
๐Ÿ’ฐAI for Finance

We Asked 5 AI Models to Build a $10K Portfolio

Five models. Same prompt. One $10,000 portfolio test. Gemini returned the most. Claude managed risk the best. Perplexity was the easiest to defend. And the disagreements between them told us more than any single answer could.

Read article โ†’
๐Ÿค–

Stop guessing. Get verified AI answers.

Talkory.ai queries GPT, Claude, Gemini, Grok, Sonar and Kimi K3 simultaneously, cross-verifies their answers, and gives you a confidence-scored consensus. Free to start.

โœ“ Free plan includedโœ“ No credit cardโœ“ Results in seconds