There is a number in this year’s enterprise AI research that should reframe how QA leaders think about their roadmaps, and it has been largely overlooked because it sits in adoption reports rather than testing ones.
Forrester and Anaconda research published in 2026 found that 88% of agent pilots never graduate to production. When leaders were asked why, the top blocker was not cost, not model quality, and not executive buy-in. It was evaluation gaps, cited by 64% of leaders, ahead of governance friction at 57% and model reliability at 51%.
Evaluation is testing. Which means the single largest obstacle standing between enterprise AI agents and production deployment is the inability to validate that those agents work correctly.
The bottleneck for agentic AI in 2026 isn’t building agents. It’s proving they behave.
This piece pulls together what the 2026 research says about where agentic adoption actually stands, why the gap between deployment and production is so wide, and what that means specifically for ERP, the environment where agent decisions carry financial reporting consequences and the validation bar is highest.
The Adoption Numbers Tell Two Different Stories
Read the 2026 agentic AI research quickly and you get a picture of near-universal adoption. Read it carefully and you get something more interesting: a very large gap between agents being present and agents being trusted.
| 80% of enterprise apps shipped/updated in Q1 2026 embed at least one AI agent (Gartner) | 31% of organizations have an agent actually running in production (S&P Global) | 88% of agent pilots never reach production (Forrester/Anaconda, 2026) | <10% have scaled agents to tangible value in any single function (McKinsey, 2026) |
Those four figures describe the same market. Vendors have shipped agents into nearly every enterprise application. Roughly a third of organizations have one genuinely running in production. Nearly nine in ten pilots die before they get there. And fewer than one in ten have scaled anything to measurable value.
Service Now’s Enterprise AI Maturity Index adds a useful corroboration: 59% of enterprises report “using agentic AI,” but only 9% report meaningful progress on autonomous multistep workflows, which is the thing agents are actually for.
The direction of travel is not in doubt. KPMG’s Q1 2026 AI Pulse Survey of leaders at billion-dollar-plus US organizations found 54% actively deploying AI agents, up from 11% a year earlier. Gartner forecasts 40% of enterprise applications will feature task-specific agents by the end of 2026, up from under 5% in 2025. Adoption is real and fast. Production maturity is neither.
The Blocker Is Evaluation, Not Capability
The most common explanation offered for the pilot-to-production gap is that models aren’t good enough yet. The data doesn’t support that.
In the Forrester/Anaconda findings, model reliability ranks third among blockers at 51%. Evaluation gaps rank first at 64%. The distinction matters enormously: model reliability is a question of whether the agent can perform the task. Evaluation is the question of whether you can demonstrate that it did. Organizations are stalling not because their agents don’t work, but because they cannot prove it to the people who have to sign off.
This reframes agentic QA from a downstream concern into the gating function for the entire agentic AI programme. If 64% of leaders identify validation as their primary obstacle, then the testing capability isn’t something you add after agents are deployed, it is the thing that determines whether they get deployed at all.
Gartner’s forecast that more than 40% of agentic AI projects will be cancelled by the end of 2027 points to the same conclusion from the other direction. The drivers it names are escalating costs, unclear business value, and inadequate risk controls. Two of those three are downstream of the same problem: without validation you cannot demonstrate value, and without validation you cannot evidence controls.
The Governance Gap Behind the Evaluation Gap
The governance numbers explain why evaluation has become the constraint.
Research published in June 2026 found that while enterprise agentic adoption reached 72% in production contexts, 60% of organizations lack formal governance for it. A separate figure puts organizations with a mature governance model at just 21%. Both describe the same condition: agents are operating in environments where nobody has defined what “correct” means, who owns the outcome, or what evidence would demonstrate compliance.
That is a tolerable position for an agent summarizing meeting notes. It is not tolerable for an agent coding a vendor invoice or routing a procurement approval, which is precisely where ERP deployments have gone.
What’s Specific to ERP
ERP is where the general agentic adoption story stops being general, for three reasons.
The agents arrived through the vendor, not a procurement decision. SAP has embedded Joule agents across S/4HANA, Success Factors, Concur and Ariba. Microsoft has embedded Copilot agents across Dynamics 365 Finance, Supply Chain and Sales. Most organizations did not evaluate, select, or approve these agents, they arrived in a release wave. That inverts the normal governance sequence, where controls are designed before deployment.
The consequences are financial, not operational. An agent that mis-summarizes a document wastes someone’s time. An agent that codes an invoice to the wrong GL account, mis-routes an approval past a threshold, or proposes an incorrect period-end adjustment produces a financial reporting error, which carries audit, disclosure, and potentially material weakness implications.
The regulatory environment tightened in 2026. EU AI Act obligations for high-risk systems became fully enforceable in August 2026, and the SEC announced a dedicated SOX enforcement group in March 2026. Neither is AI-specific to ERP, but both land hardest on AI participating in financial processes.
Forrester’s prediction that half of enterprise ERP vendors would launch autonomous governance modules during 2026, combining explainable AI, automated audit trails, and real-time compliance monitoring, reflects vendors recognizing the same gap. The relevant question for buyers is whether a vendor-supplied governance module covering that vendor’s own agents constitutes independent validation.
How to Read Agentic Adoption Statistics
A note on methodology, because the spread in published figures is wide enough to be confusing and the confusion is being used carelessly in a lot of vendor material.
KPMG’s quarterly pulse showed agent deployment falling from 42% to 26% in late 2025 before rebounding to 54%, a swing KPMG attributes not to abandonment but to leaders adopting stricter definitions of what counts as a true agent. That single data point is a useful warning: much of the variance between surveys is definitional, not behavioral.
Three questions worth applying to any agentic adoption figure, including the ones in this article:
- What is actually being measured? “Using AI agents” can mean anything from a pilot in one team to production deployment across a business unit. The definition usually sits in a footnote.
- Who was surveyed? Billion-dollar-plus US enterprises produce very different numbers from a global mid-market sample. KPMG’s 54% describes large US organizations, not the market as a whole.
- Does it reflect usage, rollout, or outcome? These are three different things, and the gap between them is the entire story of 2026. An 80% embedding rate and a sub-10% scaled-value rate are both true simultaneously.
Trust the direction of travel more than any single decimal. The direction is unambiguous: agents are being deployed far faster than they are being validated.
What This Means Going Into 2027
Four things follow from the data with reasonable confidence.
- Validation capability becomes a procurement criterion. If evaluation gaps are the top production blocker, organizations will start asking agent vendors how their outputs can be validated independently, and treating an unsatisfying answer as a reason not to buy.
- The cancellation wave Gartner forecasts will disproportionately hit ungoverned deployments. Inadequate risk controls is one of three named drivers behind the projected 40%+ cancellation rate by end of 2027. Deployments with demonstrable validation are structurally better positioned to survive a budget review.
- Vendor-supplied governance modules will not fully satisfy auditors. Forrester expects half of ERP vendors to ship autonomous governance modules in 2026. Independent validation of a vendor’s own agents is a different assurance proposition than the vendor’s own monitoring, and audit practice tends to recognize that distinction.
- Multi-agent deployments raise the validation bar again. 22% of production deployments now coordinate three or more agents. Validating a single agent’s decisions is tractable; validating the emergent behavior of several interacting agents across ERP modules is a materially harder problem that most organizations have not yet encountered.
The Testing Layer Is the Gating Function
The through-line across every figure in this analysis is the same. Agents are shipping faster than organizations can validate them, and validation, not capability, not cost, not model quality, is what the people responsible for these deployments name as their primary constraint.
For ERP specifically, that constraint is sharpest, because agent decisions there produce financial reporting outcomes subject to external audit. The organizations that get agents into production and keep them there through a budget cycle will be the ones that can demonstrate, transaction by transaction, that their agents behaved correctly and stayed within scope.
Sofy’s ERP test agents were built for that specific problem, validating business outcomes at the data layer across SAP and Dynamics 365, and producing audit-grade evidence for each validation. The five-dimension framework behind that approach covers decision correctness, scope compliance, guardrail enforcement, audit traceability, and behavioral drift.
Frequently Asked Questions
What percentage of AI agent pilots reach production?
Forrester and Anaconda 2026 research found 88% of agent pilots fail to graduate to production. The top blockers cited by leaders were evaluation gaps (64%), governance friction (57%), and model reliability (51%), meaning the primary obstacle is the inability to validate agent behavior rather than agent capability itself.
How many enterprises are actually running AI agents in production?
Roughly 31% of organizations have an agent running in production according to S&P Global Market Intelligence, despite 80% of enterprise applications shipped or updated in Q1 2026 embedding at least one agent per Gartner. ServiceNow’s Enterprise AI Maturity Index found 59% report using agentic AI but only 9% report meaningful progress on autonomous multistep workflows.
What is agentic QA?
Agentic QA covers two related things: testing performed by AI agents, and the validation of AI agents operating inside business systems. In an ERP context the second meaning dominates, verifying that autonomous agents making decisions inside SAP or Dynamics 365 produce correct outcomes, stay within authorized scope, and generate evidence an auditor can evaluate.
Why are agentic AI projects being cancelled?
Gartner forecasts more than 40% of agentic AI projects will be cancelled by the end of 2027, naming escalating costs, unclear business value, and inadequate risk controls as the primary drivers. Two of those three relate to validation, organizations struggle to demonstrate value or evidence controls without a way to verify agent behavior.
Do ERP vendors’ own governance tools solve this?
Forrester expects half of enterprise ERP vendors to launch autonomous governance modules during 2026, combining explainable AI, automated audit trails, and compliance monitoring. These help, but a vendor monitoring its own agents is a different assurance proposition from independent validation, a distinction external auditors generally recognize.
Sources
| Figure | Source |
| 88% of pilots don’t reach production; evaluation gaps 64%, governance 57%, reliability 51% | Forrester / Anaconda, 2026 |
| 80% of enterprise apps embed ≥1 agent (Q1 2026); 40% of apps with task-specific agents by end 2026 | Gartner |
| 31% of organizations have an agent in production | S&P Global Market Intelligence |
| 59% using agentic AI; 9% meaningful multistep progress | ServiceNow Enterprise AI Maturity Index |
| 54% actively deploying agents, up from 11% YoY | KPMG Q1 2026 AI Pulse Survey (March 2026) |
| ~2/3 experimented; <10% scaled to tangible value | McKinsey, 2026 |
| 72% production adoption; 60% lack formal governance | Agentic AI Institute analysis, June 2026 |
| 21% have a mature governance model; >40% of projects cancelled by end 2027 | Gartner |
| Half of ERP vendors to launch autonomous governance modules in 2026 | Forrester |
| 22% of production deployments coordinate 3+ agents | 2026 enterprise deployment analysis |
See What Validated Agent Coverage Looks Like
Watch Sofy validate agent-driven transactions across SAP or Dynamics 365, with audit-grade evidence for every decision.