Agentic AI projects in procurement fail for three reasons far more often than for technical ones: the scope is too broad, there is no measured baseline to prove value against, and the autonomy level was chosen for the demonstration rather than for the reversibility of the work. Gartner predicts that more than 40% of agentic AI projects will be scrapped by 2027, citing legacy systems and unclear cost, while separately forecasting a growing market; both are what an early technology looks like.
Why This Matters
The gap between starting and scaling is wide. The Hackett Group found only 12% of procurement organizations running AI at large scale in 2026, and EY found 80% of CPOs planning to deploy generative AI within three years while 36% had meaningful implementations. In a July 2026 BCG study of more than 200 procurement and technology leaders, nearly half cited difficulty integrating agents with ERP and procure-to-pay systems, about a third cited inconsistent data quality, and 71% cited trust. None of those is a statement that the technology does not work; they describe readiness.
How It Works
Scope too broad. "Automate procurement" is not a project. "Let the agent triage inbound quote replies for one category" is. Broad pilots produce broad, unmeasurable results and are cancelled at the first budget review because nobody can say what improved.
No baseline. A pilot without minutes per request, response rate, rework and cycle time captured beforehand cannot demonstrate a gain. When the first mistake happens, the argument is anecdote against anecdote, and caution wins.
Autonomy chosen for the demo. Full autopilot on day one impresses a sponsor and produces the incident that ends the programme. Autonomy should follow reversibility: research first, then drafting, then sending under a template, with the award staying human throughout.
Data and integration underneath. The organizational findings point the same way. The strongest outcomes in BCG's study came from redesigning the process before deploying the technology, building capability alongside it, and establishing data and integration foundations first. Teams that skip to the interesting stage inherit the failure modes of the stages they skipped.
The sequence that avoids these, six stages with an evidence gate at the end of each, is in the AI transformation roadmap for procurement; the guardrails are in what to let an AI agent do in procurement.
FAQ
Is the 40% scrapped figure a reason to wait?
No. It is a reason to scope narrowly. The same forecaster expects supply-chain software with agentic AI to reach $53 billion in spend by 2030 and 40% of procurement teams to run at least one agent by 2028. A high project failure rate inside a growing market is normal for early technology, and the failures are concentrated in avoidable causes.
Which task should a first pilot use?
The one you do most often whose mistakes are cheapest to undo. Quote extraction and normalization fits most teams, because it structures inbound data rather than depending on your own records being clean, and a wrong extraction is caught at comparison rather than at a supplier.
How do you tell a failing pilot from a slow one?
A slow pilot shows falling escalation rates and stable response rates against its baseline. A failing pilot has no baseline, a scope nobody can state in one sentence, or a rising exception rate that nobody has agreed a threshold for.
People also search for:
