enterprise-ai

Where's the ROI? The Question Your CFO Can't Answer

Magnus Hedemark 6 min read
Vintage steel engraving of a Black man in his 50s in a dark blazer, hand resting on an empty boardroom chair, long table receding toward dusk-lit windows
The question hangs in the room: is the AI working? The answer cannot be silence forever.

The scene is always the same. The budget review runs long, the room gets quiet, and someone asks the question nobody has rehearsed: "Is the AI working?"

The silence is the answer. Not because the executive is unprepared. Because the question is unanswerable at most companies. Not because AI fails. Because nobody built the measurement layer that would let anyone prove it works.

The numbers make the silence legible. A Deloitte survey of 1,326 global finance leaders found 63% have fully deployed AI in their finance function. Only 21% say those investments deliver clear, measurable value. Just 14% have integrated agents. The gap between adoption and demonstrated value is roughly three to one, and it is not a finance problem. It is the defining condition of enterprise AI in 2026.

The answer is silence because the question was never answerable

There is a name for what happens when an organization cannot tell whether AI caused an outcome. The AI Evaluability Gap research from Srivastava and Sah calls it governance non-identifiability: outcomes alone cannot warrant decisions. The practical translation is brutal. You cannot tell whether the AI did the work, whether the background market did it, or whether the numbers were always going to move that way.

Vintage steel engraving diagram showing adoption running ahead of demonstrated value: 63% deploy, 21% report measurable value, 14% integrated agents
Adoption runs roughly three times ahead of demonstrated value in enterprise AI

The MIT GenAI Divide study, which Fortune covered, reviewed 300 public deployments and surveyed 153 executives. It found 95% of generative AI pilots delivered no measurable profit-and-loss impact. The finding is contested, which is worth saying plainly: a Marketing AI Institute critique argues the framing overreaches. But the underlying diagnosis survives the critique. Most AI projects are approved on projected ROI that nobody ever goes back to validate. This is the pilot-failure pattern Groktopus examined in why AI code pilots die: the tools work, the process around them does not.

That is the smoking gun. MIT Sloan's analysis found 61% of enterprise AI projects were approved on projected ROI that was never measured again after launch. And 73% of failed AI projects had no agreed definition of success before the work began. The ROI was never unanswerable because the technology failed. It was unanswerable because nobody defined what a good answer would look like, then never checked whether the projection matched reality.

Every enterprise is running on faith, and most know it

This is not a marginal finding. McKinsey's State of AI research shows more than 80% of organizations report no tangible EBIT impact from generative AI, even as 88% experiment with it. The HBR Analytic Services survey sponsored by Appian, which polled 385 decision makers in March 2026, found only 16% report realizing a high degree of measurable value. Eight percent report no measurable value at all.

Vintage steel engraving two-panel diagram: AI alongside work shows 18% embedded and 34% standalone, AI embedded in workflows shows 71% seeing substantial value
Value concentrates where AI is embedded in workflows, not where it runs alongside them

The gap is not that AI fails. The gap is that value concentrates where measurement exists, and evaporates where it does not. This is the same pattern Groktopus documented in how AI amplifies what already exists: the multiplier works for the organizations that built the foundations, and works against the ones that did not. The same HBR survey found 71% of organizations that embed AI into workflows see substantial or moderate value, while only 18% have AI primarily integrated into how work happens. Adoption is a leading indicator. Value is a lagging one. Most organizations never built the bridge between them.

Finance is where this becomes undeniable, because the CFO cannot fake the answer. The finance function reconciles to the penny and answers to regulators. It cannot wave through unverifiable automation on the strength of a demo. That is why the Deloitte finance-trends data is the cleanest evidence of the pattern: the function with the highest verification burden is the one where the gap between deployment and demonstrated value is most visible.

The "give it time" argument is real, and it is not a free pass

The honest counterargument is that AI value follows a J-curve: a dip while teams learn, then a rebound. The evidence supports the dip. It does not support treating the dip as an excuse.

Vintage steel engraving J-curve diagram showing a 1.33 percentage point initial drop, four-year recovery, and older firms losing about a third to declining management practices
The dip is real, but it is steeper and longer without management discipline

Erik Brynjolfsson, Kristina McElheran, and colleagues analyzed two Census Bureau surveys covering tens of thousands of manufacturing firms. They found AI adoption initially drops productivity by 1.33 percentage points, a figure that jumps to roughly 60 points when corrected for selection bias, then recovers over a four-year window. The firms that recover fastest are the ones already digitally mature. The firms that struggle are the older ones, and the research shows that declining management practices account for nearly a third of their losses.

Read that carefully. The J-curve dip is not a neutral tax everyone pays. It is steeper and longer when the organization lacks the data foundation, the management discipline, and the measurement infrastructure to flatten it. The dip is tuition, but the tuition buys nothing unless the organization is learning the right practices while it pays.

Even the canonical J-curve number is softer than it looks. The DORA 2026 report on the ROI of AI-assisted software development uses a 15% productivity drop over three months as the default in its sample ROI calculator, but it is explicit that this is a placeholder, not a measurement. The 2025 DORA findings are sharper: AI adoption increases throughput and instability together, which is exactly what a missing verification layer predicts. More output, more risk, no way to tell which is which.

The way out exists, and a bank proved it

DBS Bank did not run on faith. In 2022 it publicly set an ambition to generate SGD 1 billion in AI economic value within five years. In its FY2025 annual report, it disclosed hitting that mark: more than 2,000 models across 430-plus use cases. The number is defensible because the method is defensible. DBS uses a control-group benchmarking approach, comparing customer outcomes from AI-powered solutions against a matched control group. The SGD 1 billion is the measured lift, not gross revenue flowing through an AI surface.

Vintage steel engraving four-stage pipeline: pre-registered target, measured baseline, control group, number that survives audit, with DBS's SGD 1 billion measured lift noted
DBS proved the way out: a pre-registered target, a baseline, and a control group

That is the whole difference. DBS did not answer "is the AI working?" with a demo or an anecdote. It answered with a pre-registered target, a baseline, and a control group. The same discipline is available to any organization willing to build it.

The practitioner frameworks converge on the same shape. Define success before you build. Measure the baseline before you deploy. Track unit cost reduction, revenue lift, and risk avoidance separately, because each lands on a different line of the P&L. And make the CFO able to reproduce the measurement, because a number nobody can recompute is not a number, it is an anecdote in a spreadsheet.

The question every enterprise must answer

The board is not going to stop asking "is the AI working?" The question only gets louder as AI spend grows. And the answer cannot be silence forever.

Vintage steel engraving process diagram from hope through the measurement layer to a claim and a defensible number
The measurement layer turns AI spend from a hope into a claim the CFO can defend

The fix is not a better model. It is a measurement layer: a pre-registered hypothesis, a baseline, a control mechanism, and a number that survives audit. Without it, every AI dollar is a bet that the organization is the 21% that can show value rather than the 63% that deployed and hoped. And the pilot numbers most organizations approved on were never honest to begin with, as Groktopus showed in why pilot economics are lies: costs that look irresistible in the demo multiply at production scale, while the value side goes unmeasured.

Finance is the canary because it cannot fake the answer. But the measurement layer is not finance's problem to solve alone. It is the shared infrastructure that turns AI spend from a hope into a claim, and from a claim into something the CFO can defend in the room where the silence used to live.