Why AI pilots stall before they show ROI

Most AI pilots die at the demo because they target applause, not a metric. Here's how the ROI of AI gets lost and how to design a pilot that proves it.

Shahriar P. ShuvoShahriar P. ShuvoAI ROI & Strategy7 min read
Why AI pilots stall before they show ROI

You ran the pilot. It demoed beautifully. The room nodded, someone said "ship it," and three months later the feature is either still in staging or live and ignored. Nobody can say what it earned. That is the moment the roi of ai quietly disappears, and it happens to good teams constantly.

The cause is almost never the model. It's that the pilot was built to win a demo, not to move a number your business already tracks. A demo has no baseline and no owner, so when finance asks "what did it return?" there's no answer. Unanswered AI work gets defunded.

This piece is about that gap: why pilots stall at the proof-of-concept line, and how to design one backward from a single metric so it produces an ROI figure you can defend.

Where the roi of ai disappears when a pilot stalls

A pilot stalls when nobody can name the metric it moved. That's the whole diagnosis. The technology usually works; the case for keeping it doesn't exist because it was never built.

The pattern is well documented. Gartner predicts that at least 30% of generative AI projects will be abandoned after proof of concept by the end of 2025, citing unclear business value alongside data and cost issues. Their own analyst put it plainly: organizations are "struggling to prove and realize value." Proof is the bottleneck, not capability.

BCG's data tells the same story from a different angle. In its 2024 research, nearly half of companies were still parked at the proof-of-concept stage (49%), with only 22% scaling value and 4% running mature AI value engines. The pilot line is where most AI work goes to sit.

The split between the pilots that cross that line and the ones that don't comes down to one design choice made at the start. If you want the full version of this, our cornerstone on how to tie one AI feature to one metric you already track walks the whole method. Here we'll stay on the pilot.

WARNING

A demo proves the feature can work. It proves nothing about whether it should exist. If your pilot's success criterion is "people in the meeting were impressed," it has no ROI and it will stall.

Why do AI pilots fail to scale

Pilots fail to scale because three things are missing, and each one independently makes the ROI number unprovable.

No baseline. You can't show a lift if you never recorded the starting point. Most pilots skip this because the team is racing to a demo date, and a baseline feels like overhead. It isn't. It's the only thing that turns "people seem to like it" into a number.

No counterfactual. Without a holdout or a before/after on the same cohort, any change in the metric could be seasonality, a pricing tweak, or a good month. You measured movement, but you can't attribute it to the AI.

No owner of the metric. A pilot owned by "the AI team" optimizes for a working model. A pilot owned by whoever is accountable for retention or activation optimizes for the number. Different owner, different pilot.

When those three are absent, scaling is a leap of faith, and most leaders won't take it. BCG found that only 26% of companies have built the capabilities to move beyond proofs of concept and generate tangible value, across a survey of 1,000 executives. The other 74% aren't short on models. They're short on proof.

Why does my AI pilot show no ROI

Because "show" requires a before and an after on the same metric, and your pilot almost certainly never captured the before.

ROI is a delta. Delta needs two measurements. If the pilot launched without a recorded baseline, you have one measurement and a feeling. That's why measuring ai roi after the fact is so painful: you're reconstructing a number that was never set up to be measured. The fix is upstream, before you ship anything, which is why we treat setting a baseline before you ship as step zero, not a nice-to-have.

Here's the difference in how the two kinds of pilot are scoped:

QuestionDemo-scoped pilotMetric-scoped pilot
What does success look like?"The room is impressed"The chosen metric moves past a defined threshold
Baseline captured first?NoYes, before launch
Who owns the number?The AI teamThe person accountable for that metric
Counterfactual?NoneHoldout or before/after on one cohort
Output at the endA working featureA defensible ROI figure (or a clean "kill it")
What happens nextStalls, awaiting "value"Scales, kills, or iterates on evidence

The demo-scoped column is what most teams run. It produces something that works and proves nothing. The metric-scoped column produces a smaller, less flashy pilot that answers the only question that matters to a budget holder.

How to get an AI pilot past the demo

Design the pilot backward from one metric you already track. Pick the metric first, set its baseline, then choose the smallest AI feature that could plausibly move it. The model comes last, not first.

That ordering is what gets a pilot past the demo, because it builds the ROI argument into the design instead of bolting it on after. It also lets you project the ROI before you write a line of code, so you know whether the pilot is worth running at all.

The 5-step metric-first pilot (run in this order)
 
1. PICK ONE METRIC you already track and report.
   (retention, activation, conversion, expansion. one, not three)
 
2. SET THE BASELINE before anything ships.
   Record the current number + how it's measured. This is your "before."
 
3. SCOPE THE SMALLEST FEATURE that could move that one metric.
   Not the most impressive AI. The least AI that still moves the number.
 
4. DEFINE THE THRESHOLD + KILL CRITERIA up front.
   "Activation +X points in N weeks, or we shut it off." Decide before you're attached.
 
5. RUN WITH A COUNTERFACTUAL.
   Holdout group or clean before/after on one cohort. No counterfactual, no ROI.
 
Result: a number you can defend, or a clean reason to kill it. Both are wins.

Scoping the smallest feature in step 3 is how you de-risk the pilot. A narrow pilot that can still move a metric costs less, ships faster, and gives you a clean read. A sprawling pilot that tries to be impressive is slow, expensive, and muddy on attribution. Gartner notes the same financial trap: as GenAI initiatives widen, the cost of building and deploying them climbs, which is part of why so many get cut. Smaller is not just safer. It's a better experiment.

What a pilot that protects ROI on AI investments looks like

A pilot that protects roi on ai investments is the first proof point in an ai adoption strategy, not a science fair entry. It exists to answer one question with evidence: does this AI feature move a metric we care about, enough to justify building it for real?

That means it has an owner accountable for the number, a baseline captured before launch, a threshold and kill criteria agreed in advance, and a counterfactual so the result is attributable. When it ends, you either have a defensible figure or a clean reason to walk away. Both outcomes protect your budget. The only bad outcome is the pilot that ran, demoed, and left you with a feeling.

This is also where honesty earns trust. We frame these results as projected, not achieved: a Concept Demo built to show a metric is designed to move, with the assumptions on the table. A projected number with visible assumptions beats a vanity demo with none, because it's the kind of number you can actually defend to leadership when they ask what it returned.

The teams that do this pull away from the ones that don't, and the distance is growing. A year after its first study, BCG found that the gap between the few companies generating real value from AI and everyone else only widened, with a small group capturing outsized returns while most stayed stuck. The differentiator was never model access. It was the discipline to make every pilot move a metric and prove it.

If your last pilot stalled, the next one doesn't have to. Start it from a metric you already track, set the baseline before you ship, and decide what would make you kill it. Do that and the roi of ai stops being a story you tell a board and becomes a number you can stand behind.

TIP

Want to know which AI feature would actually move a metric you already track, with a projected ROI before you build? How the AX Audit works.

AI Experience (AX) Audit

Find out which opportunity is actually worth building

The audit looks at your product and your metrics, then tells you where AI earns its place and where it does not.