How to measure AI ROI: set a baseline first

How to measure AI ROI starts before you ship. Capture a metric baseline first, so the after-number is provable evidence and not a story you tell later.

Sohanur RahmanSohanur RahmanAI ROI & Strategy7 min read
How to measure AI ROI: set a baseline first

A team ships an AI feature. A month later churn is down two points and someone asks the only question that matters: did the AI do that, or did the quarter just go well? Nobody in the room can answer. The numbers moved, but the cause is a guess.

That gap is the whole problem with how to measure AI ROI, and it almost always traces back to one missing step. Nobody wrote down the "before." Without a number captured while the AI was off, every after-number is a story, not a measurement. You can describe what happened. You cannot prove what caused it.

The fix is unglamorous and cheap, and it has to happen before the model ships. This post is about that step: capturing a metric baseline so the after-number actually means something.

Why a baseline is the whole answer to how to measure AI ROI

You cannot measure a delta without a starting point. A baseline is the value of one metric you already track, recorded before the AI goes live, so that any later change has something to be measured against. No baseline, no ROI. Everything else is commentary.

The reason this matters now is that access to AI is no longer the bottleneck. 78% of organizations reported using AI in 2024, up from 55% the year before, according to Stanford's 2025 AI Index. Shipping the feature is the easy part. Proving it paid off is where teams fall down, which is the same wall the cornerstone on measuring the ROI of an AI feature runs into: the method only works if step zero, the baseline, exists.

The cost of skipping it shows up at scale.

MIT's 2025 study, The GenAI Divide, found that roughly 95% of enterprise generative AI pilots produced no measurable impact on profit and loss. The word that does the work there is measurable. Plenty of those pilots may have helped. Almost none could prove it.

That is the trap. When measuring AI ROI is an afterthought, you end up in the 95%: a feature that might be working and a number you cannot defend in a board meeting.

Why does a baseline matter for AI ROI

Because without one, you cannot separate the AI from everything else that changed. A baseline matters because business metrics move for a dozen reasons at once, and attribution is impossible after the fact if you never isolated the variable.

Three failure modes show up every time:

  • The unprovable delta. Churn dropped. So did the price of a competitor, and you also shipped a pricing page redesign the same week. Which one moved the metric? Without a baseline and a control, you are guessing.
  • Regression to the mean. You shipped the AI during a bad month. The next month was always going to be better. The recovery gets credited to the AI that did nothing.
  • Seasonality. Activation always rises in January. Ship an onboarding copilot in December and the seasonal bump looks like your win.

A baseline does not fix attribution on its own, but it is the precondition for everything that does. It turns "the number went up" into "the number went up from this specific point, measured this specific way," which is the first thing anyone reviewing your ai roi metrics will ask for. Each failure mode above is one of the traps that quietly inflate a result, and measuring AI ROI honestly means catching them on purpose instead of letting a flattering number stand.

How to set a baseline for AI ROI in five steps

Setting a baseline is a one-week job, not a quarter-long project. Here is the sequence.

  1. Pick one metric you already track. Not a new metric invented for the AI. Churn, activation rate, time-to-resolution, conversion, expansion. The point of an AI feature is to move a metric the business already cares about, so measure against that.
  2. Define the window and the cohort. State exactly what you are measuring: 30-day logo churn for accounts created in the last two quarters, for example. Ambiguity here is where after-numbers go to die.
  3. Capture the pre-AI value. Record the metric with the AI off, across a window long enough to smooth out noise. One bad week is not a baseline.
  4. Hold out a control. Ship the AI to part of your users and keep a comparable group without it, or stagger the rollout by segment or geography. The control is what lets you attribute the change to the AI rather than to the calendar. Nielsen Norman Group's productivity research got its 66% figure precisely because it measured users on the same tasks with and without AI; the control group is the whole reason the number is trustworthy.
  5. Write down the threshold. Before launch, state the number that would count as success and the number that would count as a flop. This is where you run the projection before you build, so the baseline feeds a real target instead of a hopeful one.

Here is the shape of the artifact you are producing:

# baseline.yaml (captured BEFORE the AI ships)
metric: "30-day logo churn"
cohort: "accounts created in last 2 quarters"
window: "8 weeks pre-launch"
baseline_value: 4.2%        # measured, AI off
control_group: "50% holdout, randomized at account level"
success_threshold: "<= 3.4%"   # the after-number that earns the build
flop_threshold: ">= 4.0%"      # kill or rework if we land here

And the formula the baseline feeds, which is useless without that baseline_value:

AI ROI = (value of the metric delta - cost of the AI feature) / cost of the AI feature
 
metric delta = baseline_value - post_launch_value   (attributed via the control group)

A baseline-skipped launch and a baseline-first launch look identical on day one. They look very different the day someone asks for proof.

Baseline captured firstBaseline skipped
Before-numberRecorded, with AI offReconstructed from memory, or guessed
AttributionControl group isolates the AI"It probably helped"
The after-numberA measured deltaA story
Board question "did it work?"Answered with a figureAnswered with a shrug
Decision it supportsBuild more, kill, or holdNone you can defend

How to measure AI impact without a baseline

The honest answer is that you mostly can't, and anyone who tells you otherwise is selling a calculator. If the AI already shipped and no baseline exists, you are reconstructing, and reconstruction is weaker than measurement every time.

There are salvage options, in rough order of trustworthiness:

  • Cohort backfill. If you have historical event data, rebuild the pre-AI metric from logs. This works only if the data was clean and the definition has not drifted.
  • Staggered rollout from here. You cannot recover the past, but you can create a control going forward by rolling the AI back for a holdout segment and measuring the difference now.
  • Geography or segment holdout. Compare regions or plan tiers where the AI is on against comparable ones where it is off. Imperfect, because the groups are not randomized, but better than nothing.

Each of these is a way to de-risk a missing baseline, not a substitute for one. The takeaway is simpler to act on before launch than after: capturing the before-number costs a week. Reconstructing it costs your credibility and often is not possible at all.

WARNING

Do not let a vendor's projected-ROI spreadsheet stand in for a measured baseline. A projection with no measured starting point is fiction with decimals. The input that makes it real is the number you captured with the AI off.

Turning the baseline into an AI ROI report leaders trust

An ai roi report is just the before-number, the after-number, the control, and the delta, presented so a skeptical reader can follow the attribution. The baseline is what makes the report a measurement instead of a narrative.

What a credible report contains:

  • The metric, its definition, and the window, stated plainly.
  • The baseline value, captured pre-launch.
  • The post-launch value, with the control group's value beside it.
  • The delta, attributed to the AI via the gap between treatment and control.
  • The cost, so the delta becomes ROI and not just a nicer chart.

This is not a new idea. It is the discipline of controlled experiments applied to AI features. Decades of online experimentation show that even expert judgment is a poor substitute for a controlled comparison; the canonical example is a small Bing headline change that intuition shelved for months and a holdout test revealed was worth a 12% revenue lift. The baseline plus the holdout is what converts opinion into evidence.

TIP

Set the bar before you read the result. Decide what return you should actually expect up front, so the after-number gets judged against a threshold you committed to, not one you rationalize backward once the data is in.

The deeper point about measuring AI ROI is that the expensive mistake is invisible at the time you make it. Skipping the baseline costs nothing on launch day and everything on the day you have to defend the spend. The before-number is cheap to capture and impossible to recover. So the real answer to how to measure AI ROI is not a formula you apply afterward. It is a number you write down before the model ever sees a user.

TIP

Want a measured baseline and a projected delta on a metric you already track, before you build? How the AX Audit works.

AI Experience (AX) Audit

Find out which opportunity is actually worth building

The audit looks at your product and your metrics, then tells you where AI earns its place and where it does not.