SOX Testing Checklist for AI Agents in ERP

A practical, six-part SOX Testing Checklist for AI agents operating inside SAP and Dynamics 365. Built for the 2026 enforcement environment.

We argued previously that in an agentic ERP, continuous testing is the control layer,  the mechanism that proves to auditors that autonomous agents stayed within bounds. That was the argument. This is the tool.

What follows is a six-part checklist for bringing AI agents operating inside SAP or Dynamics 365 into SOX scope properly,  built for the enforcement environment that actually exists in the second half of 2026, not the one that existed when your current control matrix was written.

Print it, take it into your next SOX planning session, and work through it with your control owners. Every item specifies the evidence to retain and a suggested owner, because a checklist that doesn’t name evidence is just a list of good intentions.

Three developments this year moved AI agents in ERP from a governance discussion to an audit exposure:

  • SEC dedicated SOX enforcement group (March 2026). The SEC announced a group specifically targeting audit and ICFR failures, signalling materially heightened scrutiny and lower tolerance in upcoming cycles. AI-touched controls are squarely in scope.
  • EU AI Act full enforcement (August 2026). Most obligations for high-risk AI systems are now enforceable,  ongoing risk management, technical documentation, logging, and post-market monitoring. For AI influencing financial reporting or procurement approvals, auditors will expect traceable logic, data lineage, and hard evidence of what was authorised to act.
  • COSO guidance on audit trails (February 2026). COSO was explicit that prompts and configurations are part of the audit trail, not adjacent to it,  and that the resulting evidence needs to meet PCAOB AS 1105.

What has not changed is equally important. There is still no AI-specific PCAOB standard as of mid-2026,  the June 2024 technology-assisted analysis amendments explicitly excluded AI from scope, and practitioners are applying AS 2201, AS 1105, and AS 2601 by analogy. No regulator issued AI-specific guidance in the June 2026 revision of the SEC’s Financial Reporting Manual either.

AI does not change what SOX 404 requires. It changes how you satisfy those requirements,  and where you can fail.

That gap is why a checklist is more useful right now than waiting for a standard. In June 2026, FEI’s Committee on Corporate Reporting,  chief accounting officers and controllers from companies including Meta, Walmart, ServiceNow, and Alphabet,  published a framework for managing AI risk inside ICFR, built on the SEC’s existing ICFR definition and mapped to the 2013 COSO framework your SOX program already runs on. The direction of travel is clear even without a formal standard: extend what you already do, don’t invent a parallel regime.

Before any checklist item, get the scoping question right, because it’s the one most organisations answer incorrectly. Auditors are not asking where you have deployed AI. They are asking where AI participates in a process that is within ICFR scope.

The distinction matters enormously in ERP. An AI agent summarising a dashboard is not in scope. An agent coding a vendor invoice to a GL account, routing a procurement approval, or proposing a period-end adjustment is participating in a process that feeds financial reporting,  which means the controls over that agent are part of your ICFR. A control deficiency in that agent that produces a material misstatement is a potential material weakness.

You cannot control what you have not inventoried,  and most organizations discover mid-audit that AI is embedded in ERP modules they never explicitly provisioned.

 CheckEvidence to RetainOwner
Identify every AI agent or AI-enabled feature touching financial processes in your ERP,  including vendor-embedded Copilot/Joule features you did not separately procureAgent inventory registerIT / ERP owner
For each agent, document: vendor, model or feature version, deployment date, last update date, and the specific process it participates inInventory register fieldsIT / ERP owner
Classify each agent by ICFR relevance: key control, supporting control, or out of scope (efficiency only)Risk-tier classificationInternal Audit
Confirm whether the vendor provides a SOC 1 Type II report covering the AI functionality specifically,  many do not yetVendor SOC report or documented gapVendor Management
Document the control objective each in-scope agent supports, in the same language as your existing control matrixUpdated control matrix entriesSOX PMO
Re-run this inventory after every ERP release wave,  vendors add agent capabilities without customer actionDated inventory refresh logIT / ERP owner

Agents operate through service accounts and API credentials that frequently sit outside the access recertification cycle applied to human users,  and traditional SoD analysis may not flag them because the agent is not a “user” in the GRC sense.

 CheckEvidence to RetainOwner
Every agent identity is provisioned through the same lifecycle process as human identities,  request, approval, periodic recertification, deprovisioningAccess provisioning recordsIT Security
Agent authorisations are scoped to least privilege for its documented process,  not inherited from a broad service accountAuthorisation object / role assignment exportIT Security
Agent authorisation thresholds (e.g. approval limits, transaction value caps) are documented and enforced in configuration, not just policyConfiguration evidenceERP Security
SoD analysis has been extended to cover agent identities, including combinations where an agent both initiates and approvesSoD conflict report incl. agent IDsGRC / Internal Audit
Negative testing confirms the agent refuses or escalates when presented with out-of-scope transactionsTest execution evidenceQA / SOX Testing
Agent credentials are rotated and monitored on the same cadence as privileged human accountsCredential rotation logIT Security

This is the ITGC that breaks hardest under agents. Traditional change management assumes deterministic software changed through a controlled release. An agent’s behaviour can shift when its model updates, its configuration changes, or its context evolves,  none of which necessarily passes through a change advisory board.

 CheckEvidence to RetainOwner
Vendor model or agent version updates are tracked and assessed for ICFR impact before or immediately after they take effectVersion change log with impact assessmentIT / ERP owner
Prompt and configuration changes are version-controlled and retained as part of the audit trail,  per COSO’s February 2026 positionPrompt/config version historyERP owner / SOX PMO
Any change to agent scope, thresholds, or authorisations follows a documented approval pathChange request recordsChange Advisory Board
Post-change validation confirms the agent still produces correct outcomes for a defined regression set before it resumes unsupervised operationPost-change test evidenceQA / SOX Testing
Release-wave changes (SAP support packs, D365 release waves) trigger a scheduled re-validation of all in-scope agentsRelease-wave validation recordQA / ERP owner

This is where the control either exists or doesn’t. Periodic sampling was defensible when a human approved dozens of transactions; it is not defensible when an agent processes thousands per hour with variable reasoning.

 CheckEvidence to RetainOwner
Every in-scope agent decision produces a record containing: timestamp, agent identity and version, input that triggered the action, the decision made, the rule or policy basis, and the resulting outcomeDecision-level audit recordsQA / SOX Testing
Agent decisions are validated against defined business rules continuously, not sampled periodicallyContinuous validation resultsQA / SOX Testing
Validation asserts on the business outcome (correct GL account, correct approval route, balanced entry),  not merely that the transaction posted without errorOutcome-level test assertionsQA / Controllership
Exceptions and escalations are logged, routed to a named owner, and resolved within a defined SLAException register with resolutionControllership
Audit records are immutable,  retroactive editing is technically prevented, not merely prohibited by policyImmutability control evidenceIT Security
Evidence is retrievable on demand in a form an external auditor can evaluate, not reconstructed during audit seasonSample evidence packageSOX PMO

Drift is the exposure most SOX programs have no control for, because it produces no errors. An agent whose coding accuracy declines from 94% to 87% throws no exceptions and fails no postings,  it just slowly degrades the ledger.

 CheckEvidence to RetainOwner
A behavioural baseline was captured at deployment: the expected distribution of agent decisions across accounts, routes, or categoriesDocumented baselineQA / Controllership
Actual decision distributions are monitored against that baseline on a defined cadenceDrift monitoring reportsQA / Internal Audit
Variance thresholds are defined, approved, and trigger investigation when breachedThreshold documentation + alertsInternal Audit
Baselines are re-established and re-approved after any legitimate change in business process or agent scopeRe-baselining recordsSOX PMO
Drift findings feed into the deficiency evaluation process like any other control exceptionDeficiency log entriesInternal Audit

Big Four firms have trained audit staff specifically to scrutinize AI-touched controls, AI-generated evidence, and AI-driven exception resolution. Assume the walkthrough will be detailed.

 CheckEvidence to RetainOwner
A named human owner is documented for each in-scope agent,  accountability does not transfer to the agentRACI / ownership matrixSOX PMO
The control description in your matrix accurately describes what the agent does and what the human review layer coversUpdated control descriptionsSOX PMO
You can demonstrate, end to end, how a single agent-processed transaction was decided, validated, and evidencedWalkthrough sample packageControllership
Management review controls over agent output are documented with evidence of the review actually occurring,  not just its existenceReview sign-off evidenceControllership
The audit committee has been briefed on which agents operate in financial processes and what controls cover themAudit committee materialsCFO / Internal Audit
Alignment to a recognised AI risk framework (e.g. NIST AI RMF) is documented as a governance artefactFramework mapping documentInternal Audit

Your SOX program already runs on the 2013 COSO framework. These checklist sections map onto it directly,  which is the point. You are extending an existing regime, not building a parallel one.

COSO PrincipleApplies ToChecklist Sections
P10,  Select and develop control activitiesDefining agent boundaries and the validation that enforces them1, 2, 4
P11,  Select and develop technology controlsAgent identity lifecycle, authorisation scoping, change control over behaviour2, 3
P13,  Use relevant, quality informationInputs to agent decisions and the completeness of resulting evidence,  flagged by KPMG’s ICFR Handbook as a hot topic specifically because of AI4, 5
P16,  Evaluate and communicate deficienciesException handling, drift findings, and escalation to the audit committee4, 5, 6

Most of Checklist 4 and all of Checklist 5 describe evidence that only continuous testing can produce. Sampling cannot establish a behavioural baseline. Quarterly walkthroughs cannot detect drift. And no policy document proves an agent stayed within its authorisation boundary,  only a validation record does.

This is the practical form of the argument made in the companion governance piece: continuous validation of ERP business outcomes is simultaneously a QA function and a controls function. The same agent run that confirms a purchase order routed correctly also produces the timestamped record that an auditor needs to see.

Sofy’s ERP test agents validate business outcomes at the data layer across SAP and Dynamics 365,  asserting that postings hit the correct accounts and dimensions, that approvals followed the delegation matrix, and that outcomes fall within approved tolerances,  with each validation producing an audit-grade record. For teams working through Checklist 4, that maps directly to the decision-record and continuous-validation items.

The five validation dimensions behind that approach,  decision, scope, guardrail, audit, and drift,  line up deliberately with Checklists 2 through 5, and are worth reading in full if you’re building the testing side of this programme.

For Finance specifically, period close deserves separate attention,  its enforced task dependencies mean an agent error early in the chain invalidates every downstream validation, which is a control design issue as much as a testing one.

Based on the questions Big Four teams are now trained to ask about AI-touched controls, prepare answers to these before your walkthrough:

  • Where does AI participate in processes within ICFR scope,  including vendor-embedded features you did not procure separately?
  • Who is the accountable human owner for each agent, and what specifically do they review?
  • How do you know the agent’s decisions are correct,  and how often do you check?
  • What prevents the agent from acting outside its authorised scope, and how has that been tested?
  • How would you detect if the agent’s behaviour changed after a model or configuration update?
  • Show me the complete evidence trail for this specific agent-processed transaction.

If any of those answers depends on “we’d need to pull that together,” that is the gap to close before the audit, not during it.

Two honest caveats. First, this checklist covers AI agents operating inside your ERP and participating in financial processes. It does not cover using AI to perform SOX compliance work itself,  a related but separate question, where the prudent 2026 position remains human-supervised assistance rather than autonomous agents performing control testing, given unresolved questions about audit evidence sufficiency under PCAOB standards.

Second, this is a practical framework, not legal or accounting advice. No AI-specific PCAOB standard exists yet, and your external auditor’s expectations should drive final control design. Use this to structure the conversation with them early rather than discovering their position during fieldwork.

Are AI agents in ERP in scope for SOX?

If an AI agent participates in a process that feeds financial reporting,  coding invoices, routing approvals, proposing adjustments,  then the controls over that agent are part of your internal controls over financial reporting. The scoping question auditors ask is not where AI is deployed, but where it participates in an in-scope process. A control deficiency in such an agent that produces a material misstatement is a potential material weakness.

Is there a PCAOB standard for AI in ICFR?

No AI-specific PCAOB standard exists as of mid-2026. The June 2024 technology-assisted analysis amendments explicitly excluded AI from scope, and no AI-specific guidance appeared in the June 2026 revision of the SEC’s Financial Reporting Manual. Practitioners are applying existing standards,  AS 2201, AS 1105, AS 2601,  by analogy.

What evidence do auditors expect for AI agent controls?

Decision-level records showing what the agent did, on what input, under which rule, with what outcome,  retained immutably and retrievable on demand. COSO’s February 2026 guidance was explicit that prompts and configurations form part of that audit trail, and the resulting evidence needs to satisfy PCAOB AS 1105.

Can periodic sampling satisfy SOX for AI agents?

It is difficult to defend. Sampling was proportionate when a human approved dozens of transactions; an agent may process thousands per hour with variable reasoning, and behavioural drift produces no errors to sample. Continuous validation is the practical way to produce sufficient evidence at that volume and variability.

How do we handle agent behavior changing after a vendor update?

Treat it as a change management event even though it did not pass through your change advisory board. Track vendor model and version updates, assess ICFR impact, re-validate against a defined regression set before the agent resumes unsupervised operation, and re-baseline drift monitoring if behaviour legitimately changed.

SOX did not change. What changed is that a meaningful share of the transactions inside your ERP are now initiated or coded by something that is not a person, and your existing ITGCs were designed on the assumption that they were.

Work the six checklists with your control owners, close the evidence gaps before your auditor finds them, and treat continuous validation as what it now is, not a QA nicety, but the mechanism that makes your agent controls provable.

Generate the Evidence Your Auditor Will Ask For

See how continuous validation of SAP and D365 business outcomes produces audit-grade records for every agent-driven transaction.

See Sofy in action. Book your demo.

We’ll show you exactly how it works for your team in 30 minutes.

Scriptless test automation—no coding or framework setup

Run tests on hundreds of real iOS and Android devices

Integrate with your CI/CD in minutes

Self-healing test that adapt as your app changes

Real-time debugging with logs, crash reports, and performance data