Why Most Agentic AI Pilots Never Make It Into Revenue Workflows
Eshaan Jain is a Senior Product Manager for telecomm clients (via Mphasis) focused on AI-driven CRM strategy and enterprise workflows.
gettyI’ve sat through a lot of agentic AI demos this year. Almost all of them work. The agent reads a contract, finds the clause, fills in the field, and escalates the exception. It looks ready.
Then it doesn’t ship, or it ships and gets pulled back within a few months. MIT’s NANDA initiative, in a report covered by Fortune in August 2025, found that 95% of enterprise generative AI pilots failed to deliver a measurable effect on profit and loss. The report described the situation as a gap in how organizations integrate these tools into existing workflows, separate from the quality of the underlying models. Among the pilots who succeeded, purchased and integrated tools worked about 67% of the time, compared with about 33% for in-house-built tools.
S&P Global Market Intelligence’s 2025 enterprise AI survey, covering more than 1,000 companies across North America and Europe, found the same pattern from a different angle. The share of businesses that scrapped most of their AI initiatives jumped to 42% in 2025, up from 17% the year before. The average organization abandoned 46% of its AI proofs of concept before they ever reached production. Executives named cost, data privacy and security risk as the top reasons.
Gartner, in June 2025, predicted that more than 40% of agentic AI projects will be canceled by the end of 2027, citing rising costs, unclear business value and weak risk controls. The same research found that of the thousands of vendors marketing “agentic AI,” only around 130 actually offer real agentic capability. The rest are existing tools with an agent label added.
When I’m evaluating a vendor pitch now, I ask what, specifically, the agent does without a human reviewing each step and what happens when it encounters a case it hasn’t seen before. Most of the difference between a real agent and a relabeled chatbot shows up in the answer to that second question.
In my experience, the blocker is almost always the data underlying the workflow: how complete it is, how consistent its format is, and how much of it lives in a person’s head rather than in a system.
At Amazon, I co-built a system that extracted clauses from supply chain contracts across a $40 billion annual portfolio, eventually reaching 95% accuracy. Most of that gain came from months of work building a structured taxonomy for clause types, so the system had a consistent target to extract into. Before that taxonomy existed, the same model produced results nobody could use, because “payment terms” meant six different things depending on which template a contract used.
That kind of work rarely shows up on a roadmap as its own line item. It usually gets folded into “data prep” and scheduled for two weeks, when the honest estimate is closer to a quarter.
Quote-to-cash workflows have the same problem, spread across more systems. Product catalogs get updated in one place and referenced from memory in another. Discount approval rules live in a policy document that’s a couple of reorgs out of date. Contract terms that should drive a renewal quote sit in a PDF attachment instead of a structured field. Pricing exceptions get approved over email and never make it back into the system that’s supposed to be the record of truth.
Ask an agent to draft a renewal quote against that environment, and it does what a new hire would do: guess, ask someone or move forward with full confidence and get it wrong. The agent fails for the same reason a new hire would. The information it needs isn’t where the task assumes it will be.
PwC’s April 2025 survey of 308 U.S. business leaders, reported by Digital Commerce 360, shows the same hesitation. Seventy-nine percent of companies were already using AI agents, but only 35% had reached broad deployment, and 68% admitted fewer than half of their employees used those agents regularly. Asked which tasks they’d trust an agent with, only 20% of executives said financial transactions, compared with 38% for data analysis. That gap maps onto what I see in CPQ: Agents get comfortable fast on tasks where the underlying data is already clean, and stall everywhere else.
BCG’s Build for the Future study from September 2025, based on more than 1,250 companies worldwide, found that only 5% of companies are generating AI value at scale, while 60% report little or no measurable value at all. When BCG asked what blocks the other 95%, the top answers weren’t about the models. Seventy-nine percent of respondents pointed to a lack of expertise managing unstructured data. Sixty-eight percent pointed to lack of access to high-quality data in the first place. Both ranked above concerns about hallucinations, model accuracy and AI-driven security risk.
Four studies, four research methods, one finding underneath: The data problem shows up before the model does.
Before I put an agent into a revenue workflow, I run one test: Could a new hire complete this task using only what’s documented in our systems, with no one to ask? If the answer is no, that documentation gap is the project to fix first. The agent comes second.
This is also where I’ve seen the biggest gains come from on the CPQ and contract sides of my work: cleaning up the underlying product, pricing and contract data first, so that automation, AI or otherwise, has something reliable to act on. That means one taxonomy for discount types instead of five, and one system of record for approval thresholds instead of a policy PDF nobody rereads. It’s unglamorous work, and it’s the prerequisite the agent needs.
An agent performs only as well as the data and documentation behind it. Get that right first, and the agent part gets a lot easier. Skip it, and no model upgrade fixes what’s underneath.
Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?

