I have spent 28 years evaluating internal controls — first at PwC and KPMG, then as a Chief Audit Executive at Fortune 500 and FTSE 100 companies, and now building AssurAI. In that time, one principle has never changed: every conclusion you put in a workpaper must be reproducible.
If your external auditor asks you to re-run a Management Review Control evaluation six months later, the answer must be identical. Not similar. Identical. That is what PCAOB AS 2201 demands, and it is what Big 4 inspectors look for when they pull workpapers.
Most AI systems cannot meet this bar — not because they lack intelligence, but because they are deliberately designed to be creative and variable. That works brilliantly for writing marketing copy or brainstorming ideas. It is disqualifying for SOX compliance.
The Problem With Creative AI in Audit
Large language models have a parameter called temperature. At higher temperatures, the model introduces randomness into its outputs — the same input can produce subtly different responses each time. This is what makes ChatGPT feel conversational and engaging. It is also what makes it unsuitable for PCAOB-defensible control evaluations.
Consider what happens when you use a high-temperature AI to evaluate an MRC:
- Run 1: "Precision is adequate — the threshold of ±5% is appropriate for this account balance."
- Run 2: "Precision is partially adequate — the ±5% threshold may need refinement given the volume of journal entries."
- Run 3: "The precision of the control appears reasonable for the size of this entity."
Three runs, three materially different conclusions, none of them consistent with one another. If your external auditors pull the workpaper and re-run the same evaluation, they get a fourth answer. Your support documentation is now a liability, not an asset.
PCAOB perspective: Inspectors evaluate whether management's testing methodology is consistent, documented, and repeatable. Variable AI outputs create exactly the kind of inconsistency that triggers inspection findings.
What Deterministic AI Means
At AssurAI, every MRC evaluation runs at temperature=0. This is not a trivial configuration choice. It means:
- Given the same control description, evidence, and parameters, the AI will produce byte-for-byte identical output every single time.
- Your external auditors can re-run the evaluation independently and arrive at the same conclusion.
- Year-over-year comparisons are meaningful — differences in the evaluation reflect real changes in the control, not AI variability.
- The workpaper trail is audit-ready from day one.
This is what I mean when I say AssurAI uses deterministic AI. The model still applies sophisticated reasoning — it evaluates all five PCAOB AS 2201 MRC attributes, understands the control's purpose, and produces a nuanced assessment. But given the same inputs, it produces the same outputs. Always.
The Five MRC Attributes Under PCAOB AS 2201
AssurAI's deterministic MRC testing evaluates each control against all five attributes that PCAOB inspectors examine when assessing whether a management review control operates effectively:
| Attribute | What PCAOB looks for | What AssurAI evaluates |
|---|---|---|
| Precision | Does the control catch material misstatements at the right threshold? | Threshold vs. performance materiality; specificity of the review criteria |
| Investigation | Does management follow up on exceptions identified? | Evidence of exception resolution; escalation path documentation |
| Frequency | Is the control performed often enough to catch misstatements before period-end? | Review cadence vs. transaction volume; risk of undetected error between reviews |
| Preparer Competence | Does the person performing the review have the knowledge to identify issues? | Role description; access to relevant data; independence from the preparers |
| Documentation | Is the review documented in a way that evidences it actually occurred? | Retention of support; dated approvals; linkage to the underlying data reviewed |
Each attribute receives a structured assessment: Pass, Partial, or Fail — with a specific finding and recommended remediation. The output is formatted to match Big 4 workpaper standards, so it drops directly into your external auditor handoff package.
Year-over-Year Comparability
One of the most underappreciated benefits of deterministic AI is what it unlocks for multi-year programmes. Because the evaluation logic is identical year over year, any change in the assessment output directly reflects a change in the underlying control — not a change in how the AI felt that morning.
This matters enormously during external audit. When your auditors compare this year's MRC evaluation to last year's, they want to see a coherent narrative. "Precision improved because management tightened the threshold from ±5% to ±2%" is a defensible story. "The AI gave a different answer this year" is not.
AssurAI's Prior Year Workpaper Ingestion feature feeds last year's MRC assessments back into this year's evaluation context — so the AI can explicitly call out what changed, what improved, and what remains a gap. This is the kind of year-over-year intelligence that previously required a senior manager to spend two days cross-referencing workpapers.
Why This Matters More Than Speed
The conversation about AI in audit is often dominated by speed claims. "Cut your SOX timeline by 60%." "Generate workpapers in minutes." These are real benefits, and AssurAI delivers them. But speed is a secondary benefit.
The primary benefit is defensibility.
A fast MRC evaluation that produces different answers every time is worse than useless — it creates documentation that actively undermines your control environment. A deterministic evaluation that takes the same time as manual review but produces consistent, structured, Big 4-formatted output is transformative.
"The question isn't whether AI can evaluate an MRC faster than a senior associate. Of course it can. The question is whether the output will survive inspection. That's where deterministic AI is in a different category entirely."
The Practical Implications for Your Team
If you are evaluating AI tools for your SOX programme, here are the questions to ask:
- What temperature does the model run at? If the vendor can't answer this, assume it's not 0.
- Can you re-run an evaluation from six months ago and get the same output? Ask for a demonstration.
- Is the output formatted for Big 4 workpaper standards? "AI-generated summary" is not a workpaper. A structured five-attribute assessment with findings and remediations is.
- How does the tool handle year-over-year comparisons? If it doesn't have prior year context, every evaluation starts from zero.
These aren't academic questions. They are the questions your external auditors will ask when they review your testing documentation. Having good answers — and the workpapers to back them up — is the difference between a clean opinion and an inspection finding.
What AssurAI Produces
When you run an MRC evaluation in AssurAI, the output is:
- A structured five-attribute assessment in Big 4 workpaper format
- A pass/partial/fail conclusion for each attribute with specific supporting rationale
- Recommended remediation steps for any partial or failed attributes
- A year-over-year comparison if prior year data is available
- An overall MRC effectiveness conclusion suitable for inclusion in your management assessment
Every output is identical on re-run. Every output is traceable to the specific inputs you provided. Every output is formatted to meet the documentation standards your external auditors expect.
That is what deterministic AI looks like in practice. And in my 28 years of audit, it is the only standard that belongs in a SOX workpaper.