When a simpler feature beats an AI feature on ROI

When to use AI in product and when a form, rule, or sort wins on ROI. The cost-and-return test to run before you commit budget to an AI build.

Anamoul RoufAnamoul RoufAI ROI & Strategy7 min read
When a simpler feature beats an AI feature on ROI

There's a feature idea on your roadmap with the letters "AI" attached to it, and a board member who keeps asking when it ships. The instinct is to build the AI version because that's the one that sounds like progress. That instinct is where a lot of budget goes to die.

Here's the reframe that decides when to use AI in product: the simpler feature is the default, and the AI version has to beat it on the math before it earns the work. A form, a rule, or a sort is not the fallback you settle for. It's the baseline the AI build has to outperform on cost and on the metric you already track. Most of the time it doesn't, and the people writing the check never run the comparison that would have told them.

This is the comparison. We'll cost the AI version of a feature against the boring version of the same feature, give you the break-even math, and name the four conditions that actually flip the verdict toward AI.

When to use AI in product: the boring option is the baseline

You decide when to use AI in product by costing the AI build against the simplest form, rule, or sort that moves the same metric, then keeping whichever wins on projected ROI. That's the whole rule. The AI version doesn't get a head start for being AI.

Teams skip this step for predictable reasons. The roadmap says "AI," the competitor shipped a copilot, and a form feels too plain to put in a board update. So the cheap option that would have moved the metric never gets costed, and the expensive one ships on vibes. The pattern is the same one we walk through in how to decide which AI features to build: the decision gets made on narrative instead of arithmetic.

The bill for skipping the comparison is now well documented. 95% of enterprise generative-AI pilots showed no measurable return in MIT's 2025 study, despite tens of billions in spend. Over the same period, 42% of companies abandoned most of their AI initiatives, up from 17% a year earlier, with the average organization scrapping 46% of proof-of-concepts before they ever reached production.

Most of those scrapped projects were not killed by a bad model. They were killed because a cheaper, simpler version of the same feature would have done the job, and nobody priced it.

The lesson isn't that AI doesn't work. It's that AI is the expensive option, and expensive options need to be justified against the cheap one. That's not pessimism. It's the same discipline you'd apply to any line item.

Is AI always better than a simple feature?

No. And you don't have to take an anti-hype agency's word for it. The teams who built modern machine learning say it first.

Google's first rule of machine learning is "don't be afraid to launch a product without machine learning." Their reasoning is blunt: a heuristic gets you roughly 50% of the value you'd expect from a full ML lift, for a tiny fraction of the cost and time. If you think the model will give you a 100% boost, a simple rule already gets you halfway there. "If machine learning is not absolutely required for your product," they write, "don't use it until you have data."

That halfway-for-a-tenth-of-the-price math is the heart of the decision. A lot of features capture most of their value with a rule you can write in an afternoon. The model would add a marginal lift on top, at 10x to 50x the build cost and an ongoing maintenance bill the rule doesn't have. When the simple version captures 80% of the outcome, the AI version is competing for the last 20%, and that 20% rarely pays for itself.

The comparison that decides it: AI build vs the boring build

Put the two versions of the feature side by side. This is the AI cost benefit analysis that the typical "AI ROI" article skips, because it compares AI to doing nothing instead of comparing AI to the cheaper build of the same feature.

DimensionAI versionNon-AI version (form / rule / sort)
Build cost$20k–$50k typical for a real featureDays of existing eng time
Time to shipWeeks to monthsDays
Failure riskHallucination, wrong answers, trust damageNear zero; behavior is predictable
MaintenanceOngoing: evals, drift, model updatesEffectively none
Metric targetedSame metricSame metric
Value capturedFull lift (if it works)Often 70–90% of the lift

The number that matters is not the build cost. It's the total cost of ownership against the projected metric movement. Run it before you commit:

Projected ROI = (projected metric lift x value per unit x users affected)
                ----------------------------------------------------------
                (build cost + 12-month maintenance + risk-adjusted failure cost)
 
Decision rule:
  Compute projected ROI for BOTH versions on the SAME metric.
  Build the AI version only if its projected ROI clears the simple
  version by a margin wide enough to cover the failure risk.
  A tie goes to the boring build.

This is why the verdict so often lands on the simple feature. Context helps: even when AI works, only about 25% of AI initiatives have delivered the ROI that executives expected, per IBM's 2025 CEO study. When the boring build hits 80% of the outcome at a fraction of the cost and none of the failure risk, its projected ROI is usually higher, even though the headline number is smaller.

When does a non-AI feature win

A non-AI feature wins whenever the cheaper build captures most of the value at a fraction of the cost and risk. In practice, four patterns make the boring option the right call.

  • The rule is knowable and stable. If a person could write the logic down ("flag accounts with no login in 21 days"), a rule does it faster, cheaper, and with no failure mode. You don't need a model to learn a rule you already know.
  • The data is thin. Models need volume. With a few hundred examples, a heuristic beats a model that hasn't seen enough to generalize. Build the rule now; revisit AI when the data exists.
  • A wrong answer is expensive. If the feature touches billing, compliance, or a trust-sensitive moment, a confident wrong answer costs more than the feature earns.
  • Volume is low. If the feature runs a few hundred times a month, the engineering and maintenance load of an AI pipeline never amortizes. A sort or a filter is the better trade.

WARNING

The most expensive AI feature is the one that ships, sounds smart, and quietly damages trust with a wrong answer in a high-stakes moment. If a mistake is costly and the logic is writable, build the rule. This is the core of the AI features you should not build: the ones whose failure mode outweighs their lift.

When most of these are true, the model is theater. The form moves the metric and the model just adds cost.

Should I build AI or a rules-based feature: the four-test filter

AI earns the spend when the simple build genuinely can't reach the outcome. Four tests flip the verdict from "boring build" to "this should be AI." Run them in order; the more that hold, the stronger the case.

  1. Ambiguity. The input is messy, unstructured, or open-ended (free text, images, intent) and no fixed rule can cover it. This is where models earn their keep.
  2. Scale. The feature runs at a volume where the per-decision quality gain, multiplied across millions of events, clears the build and maintenance cost.
  3. Learning. The right answer shifts over time, and a static rule would decay. A model that updates on fresh data holds its edge where a rule would rot.
  4. Defensibility. The lift compounds with your data in a way a competitor can't copy by writing the same rule. The AI is doing something the boring build structurally can't.

If fewer than two of these hold, the full when-and-when-not test almost always points back to the simple build. This is AI feature prioritization at the single-feature level: you're scoring one idea against its own cheaper twin, not against the rest of the roadmap.

When AI does win, it wins as supportive AI. It sits as a layer on top of the product, a copilot or a classifier or a search box, with reliability guardrails and a human in the loop where a wrong answer would cost you. It is never the core engine you bet the company on. The point of view is consistent: build the AI that moves a number, and only when it beats the boring build on that number.

What a verdict looks like: projected ROI on a real call

Here's the shape of the call, framed as a Concept Demo (a prototype of our thinking, with projected numbers, never a claimed client result).

The idea: a churn-save feature for a $6M ARR SaaS. The roadmap wants an "AI retention assistant." The metric is logo churn, currently 4% monthly.

The boring build: a rules-based risk flag (no login in 21 days, support ticket unresolved, seat count dropping) that triggers a save play. Two weeks of existing eng time. Designed to recover a projected share of at-risk accounts.

The AI build: a churn-prediction model plus a generated outreach copilot. Six weeks, a model to maintain, and a failure mode where a tone-deaf message annoys a healthy account.

The rules version targets the same churn metric at roughly a tenth of the cost and none of the trust risk. On projected ROI, it wins, so it ships first. The model gets revisited only once the rules version proves the metric moves and the data has accumulated to justify the lift.

That's a verdict, not a vibe. It's the same logic our AX Audit runs across a whole roadmap, ranking every idea by projected ROI and naming the ones to skip, backed by the 3X Guarantee: we find AI worth 3x the fee, or it's free.

Knowing when to use AI in product is the same skill as knowing when not to. The discipline here isn't anti-AI. It's pro-ROI, and it's what keeps your team out of the 95% that ship something clever and move nothing. Cost the boring build first, make AI beat it on the metric you already track, and you'll build less, ship faster, and have a number to show for it.

TIP

Want the head-to-head run across your actual roadmap, with each idea costed against its simpler twin? How the AX Audit works.

AI Experience (AX) Audit

Find out which opportunity is actually worth building

The audit looks at your product and your metrics, then tells you where AI earns its place and where it does not.