Most enterprises judge AI by one number and much too early in the process. Three companies show what to measure instead.
Joanne Wright had a problem that most technology executives would recognize. As IBM's Senior Vice President of Transformation and Operations, she was being asked to prove that AI was worth what the company was spending on it. So IBM had made an unusual decision: it would become its own first customer, testing AI inside its own operations before selling the results to anyone else. The company called it Client Zero.
Even more unusual, IBM did not wait for a standalone ROI figure to justify the program. Its AI performance metrics were measured separately, function by function, across human resources, finance, supply chain, technology, and sales. Some of those early efforts would have looked unimpressive on their own. Judged by a singular number in their first year, several might have been shut down.
But they weren’t. Instead, over three years, IBM reported $4.5 billion in productivity gains across more than 155 use cases.
That raises the question every executive funding AI work is living with right now: how do you tell the difference between a project that is failing and one that simply has not paid off yet?
Most enterprises collapse AI's value into a single financial number and demand it too early, then kill initiatives that were performing fine, just badly measured. Instead, organizations should match their metrics to the maturity of the use case and track value across multiple dimensions, not just dollars, so they scale what works instead of abandoning it.
AI performance metrics are hard to pin down in part because many organizations condense a multi-year, multi-function change into one metric and evaluate it before the work has matured. This results in a struggle to directly link AI with results, and there’s growing evidence for how widespread this problem is.
Bain surveyed executives in late 2025 and found that 74% of companies rank AI among their top three priorities. Yet only 23% could tie their AI performance metrics to new revenue or reduced costs. That is a considerable distance between what companies say matters and what they can actually demonstrate.
"They're really laser-focused on measuring the wrong things. There's a fundamental misunderstanding of how to measure AI."
—Shamim Mohammad, Executive Vice President and Chief Information and Technology Officer, CarMax
The Cambridge Centre for Alternative Finance surveyed 628 organizations across 151 jurisdictions in partnership with the World Economic Forum and others. It found a similar correlation. Its 2026 report found that 55% of industry respondents struggle to measure the value of their AI deployments, rising to 76% among large financial institutions. The report also found that having longer organizational exposure to AI does not solve the problem. Even among companies already scaling AI across the organization, 60% still reported difficulty measuring its value.

Meanwhile, productivity gains are widely reported by professionals. Roughly 79% of respondents in the Cambridge Centre’s study reported positive productivity effects in technology and data functions, and 75% in back-office operations. So people can feel the work getting faster; they cannot always put a number on what that is worth.
Research from the World Economic Forum's 2026 study of organizations with documented AI results found a consistent pattern: those focusing narrowly on technology or short-term return on investment consistently struggle to scale, while organizations measuring across strategy, workforce, data, technology, and governance achieve results that hold up.

These are the top three mistakes companies make when determining AI performance metrics:
"What about quality improvements? Are there fewer errors and less rework? Look at cycle times in your operations. The key is to deliberately track these metrics from day one and systematically build metric capture into your systems so you can measure the impact."
—Swami Chandrasekaran, Global Head of AI and Data Labs, KPMG LLP
In the following case studies, we’ll see how IBM, KPMG, SAP, and Foxconn headed off these issues.
IBM’s human resources teams supported employees across more than 175 countries, while procurement relied on over 40 different systems and supply chain data sat fragmented across the organization. Given that vast scope, a single metric defined too early wouldn’t have justified the full scope of what the company wanted to attempt.
So IBM structured the work differently. Teams shipped minimum viable products roughly every two weeks, emphasizing progress over perfection. Every initiative was measured for business outcomes on its own terms rather than folded into one enterprise-wide calculation. A CEO-led steering committee kept the program accountable while individual efforts developed at their own pace.
"At IBM, we made ourselves the first client, proving that transformation at enterprise scale is not only possible but delivers significant measurable value."
—Joanne Wright, Senior Vice President of Transformation and Operations, IBM
The results accumulated in pieces across functions:
Add them together, and you reach $4.5 billion across more than 155 use cases. Take any one of them in isolation during its first six months, and the case looks thin.
The insight here is about sequencing. Rather than lowering its standards for success metrics, IBM matched the standard to the stage of the work, and it kept its AI KPIs at the level where results actually appear: the function itself.
Companies often measure how widely a tool has spread rather than what it accomplished. KPMG and SAP took the opposite approach on a project with an unglamorous name: enterprise resource planning migration.
ERP migrations are slow, complicated, and expensive. They are also exactly the kind of work where a company could claim an AI win without proving one. Instead, the teams scoped a single use case and set the AI KPIs before building anything.
SAP built an AI copilot grounded in KPMG's own documentation and best practices from previous migration projects, drawing on up to nine terabytes of content across three million documents. The tool was the opposite of general-purpose: it knew one domain deeply.
What matters here is what the teams counted. They reported three figures:
Two of those three are not productivity measures. Rework is a defect rate. Delivery speed is cycle time. Only the hour and a half per day is a pure time-saved number, and it appears alongside two measures that show what happened to the work itself. That combination is what makes the result legible to a finance team. A migration that ships faster with half the rework is a clear demonstration of AI ROI.
Some of what AI delivers does not appear on a financial statement at all. But in manufacturing, this becomes unusually clear.
Foxconn, working with Boston Consulting Group on an effort called Project Genesis, faced a challenge that has little to do with cost per unit. Decades of manufacturing know-how lived with experienced masters on the factory floor. That knowledge was not written down anywhere. When those workers retired, it left with them.
“AI can power production, people power transformation.”
—Foxconn
So Project Genesis built six AI applications aimed at capturing and codifying that expertise, moving operations from intelligence to autonomy and from experience-driven workflows to continuously learning systems.
The measured results span dimensions a purely financial scorecard would miss. Workload related to production changeovers fell by half. Time to resolve problems on the line dropped 30%. Cycle times came down roughly 10%.
Overall, Foxconn’s approach treats retained human expertise as an asset worth measuring. Only the cycle time reduction translates cleanly to cost. The other two measure work quality and workload, which is precisely the non-financial value Pitfall 3 describes.
The distance between these three companies and the 77% who cannot tie AI to revenue is what they measured and when. IBM measured at the function level and let results accumulate. KPMG and SAP measured defect rates and cycle time alongside hours. Foxconn measured workload and problem resolution. None of them emphasized a single number, and none of them stopped at time saved.
Gartner's threshold explains why so many time-saved figures never reach the income statement. Below roughly 50%, a productivity gain is real but does not convert into a headcount or cost change, and most use cases sit well below that line. This is why the measures further up Gartner's value spectrum matter more. Capital use, losses avoided, customer experience, and new product contribution carry a much greater potential connection to revenue growth than productivity does.
For governance, the practical implication is that a single ROI metric applied uniformly across a portfolio can stifle projects. It preserves initiatives that produce fast results and eliminates those whose value would otherwise compound. The Cambridge research suggests why this matters more as companies grow: at large institutions, fragmented data environments, legacy systems, and rigid governance structures make single-figure attribution nearly impossible, which is exactly where the temptation to demand one is strongest.
Return to Joanne Wright's position at the start. IBM's choice was not to abandon rigor. It was to refuse a standard of evidence that the work could not yet meet, while insisting on evidence the work could produce, i.e., function-level outcomes, measured continuously, accumulated over three years, and eventually adding up to a number large enough to answer any board.
Here is the exercise worth bringing to your next leadership meeting. Pull the list of AI initiatives currently under review for funding. For each one, ask two questions:
Most organizations cannot answer the second question, and that is what is standing between them and a defensible number.
.avif)
Measures tied to capital use, losses avoided, customer experience, and new product contribution carry a stronger connection to revenue growth than productivity does. Of the ten AI value metrics Gartner has identified, productivity sits at the lowest end of the value spectrum.
Gartner reports that organizations need a 50–70% productivity gain before headcount can be reduced, while most AI use cases deliver between 0–30%. Time saved is a real result, but below that threshold, it does not indicate a cost reduction on its own.
Organizational complexity works against attribution. Cambridge research found measurement difficulty at 76% for large firms versus 49% for small ones. Fragmented data environments, legacy systems, and rigid governance structures make it harder to isolate what one AI initiative contributed to a business outcome.
Only 23% of companies can tie their AI initiatives to new revenue or reduced costs, according to Bain's late 2025 executive survey, despite 74% ranking AI among their top three priorities. Cambridge research separately found 55% of organizations struggle to measure AI value.
AI performance metrics are the measures used to determine whether an AI deployment is delivering value. The World Economic Forum and Accenture identify four dimensions: economic value, decision effectiveness, adoption and behavioral change, and risk and sustainability.