Which AI features to build (and which to refuse)

An anti-hype catalog of the AI features that demo well and move nothing, plus a one-line test for which AI features to build and which to cut in planning.

Shahriar P. ShuvoShahriar P. ShuvoAI ROI & Strategy7 min read
Which AI features to build (and which to refuse)

The hardest part of an AI roadmap is not picking which AI features to build. It is having the spine to not build the five that everyone ships. Saying no is the position, and almost nobody takes it.

The market is loud with menus of AI features to add. It is silent on the inverse: the features that look great in a sprint review, get a round of applause, and then sit in production moving no number you actually report on. This post is that inverse list. A catalog of AI features that demo well and flop quietly, each paired with the cheaper thing that already works, and a two-line test you can apply before you spend a single sprint. If you only read one section, read the table.

Refusing a feature is a real product decision, not an absence of one. The companion to it is deciding which AI features to build on purpose, scored against a metric and a cost. This piece handles the other half: what to take off the list, and why.

Deciding which AI features to build starts with a refusal list

The build decision is a subtraction problem. You do not start from zero and add the exciting ideas. You start from your backlog and remove every AI feature that cannot name a metric and a baseline. Most of them cannot.

That sounds harsh until you look at the outcomes. Across a survey of 1,000 executives, only 26% ever move past a proof of concept into tangible value. The other three quarters are not building the wrong technology. They are building features nobody can connect to a result. The fix is upstream, in planning, before the cost lands.

Here is the entire test, and it fits on two lines:

Refusal test (apply before any AI feature enters a sprint)
  1. Name the metric this feature moves      → activation, retention, time-to-value, support cost, conversion
  2. Name the baseline it moves from         → the current number, measured, today
 
If you cannot fill both lines with real numbers, the answer is no.

If the metric is "engagement with the AI feature," that is a tell. Engagement with the feature is not a business outcome. It is the feature measuring itself. To run the test properly you have to set a baseline first, which most teams skip because the baseline is often embarrassing and the AI feature is more fun to discuss.

What AI features are not worth building

These are the ones that fail the refusal test most reliably. Each demos beautifully because a demo is a controlled environment, and each moves nothing because production is not. The pattern is always the same: an AI capability is bolted onto a job a simpler tool already does, the novelty carries the demo, and adoption dies in week three.

AI feature that demos wellThe metric it claimsWhy it moves nothingThe cheaper thing that already works
Chatbot answering your own docsSupport deflectionUsers want the answer, not a conversation; it adds a turn before the linkBetter search and a sharper FAQ
AI summary of a screen the user is already onComprehension, speedPeople do not read a summary of content in front of themA clearer layout and a real empty state
Predictive score with no explanationBetter decisionsAn unexplained number gets ignored or distrusted, so no decision changesOne sortable column the user already understands
"Ask your data" natural-language query boxSelf-serve analyticsVague questions get confident wrong answers; trust collapses on the first missThree saved views that answer the real questions
Generative onboarding that writes the setupActivation, time-to-valueUsers do not trust output they did not shape; they redo itA short guided flow with sane defaults

None of these are bad technology. They are good technology pointed at a job that did not need it. Gartner attributes a large share of abandoned generative AI projects to unclear business value, which is the polite phrase for "we never named the metric." Every row above is a different way of skipping that step.

WARNING

The demo trap: a feature that wins the sprint review is selected for how it looks in a five-minute controlled walkthrough, not for whether it changes behavior over five weeks of real use. The applause in the room is the most misleading signal you will get all quarter. Treat it as noise.

How to know an AI feature will flop before you build it

You can usually call the flop in planning. The signals are loud if you are listening for them instead of for applause. A feature is likely to flop when it cannot name an outcome, when the value depends on the user trusting an output they did not shape, or when it adds a step to a job that was already one step.

That last one is feature creep wearing an AI costume: complexity added beyond the core function, dressed up as progress. The deeper version is older than this AI wave. A feature shipped is not an outcome moved, and teams that are handed features to build rather than problems to solve tend to ship a lot of the former. AI makes this trap more expensive, because the feature costs more to run and looks more impressive while doing nothing.

Score the idea before it enters a sprint. This is lighter than a full ai feature prioritization pass and catches the obvious refusals fast:

# Flop-risk scorecard (answer honestly, sum the risk)
metric_named:        true        # false = +3 flop risk
baseline_measured:   false       # false = +2 flop risk
beats_simpler_tool:  unknown     # unknown/no = +2 flop risk
trust_required:      high        # high = +2 flop risk (users distrust unshaped output)
adds_a_step:         true        # true = +1 flop risk
 
# total >= 4  → refuse, or rescope to the simpler tool
# total 1-3   → de-risk: ship the non-AI version first, measure, then decide
# total 0     → this one earns a build slot

The scorecard is not the point. The honesty is. If you fill it in the way you would defend the feature to your boss, it lies. If you fill it in the way a skeptical finance lead would read it, it tells the truth.

When to use AI in product, and when a plain feature wins

The clean rule for when to use AI in your product: use it as supportive AI sitting above a deterministic core, never as the core itself. Supportive AI suggests, drafts, ranks, and summarizes on top of a system that still works perfectly when the model is wrong or offline. That is how you de-risk the build. The product does not depend on the model being right; it gets a little better when the model is right and loses nothing when it is not.

A plain feature wins whenever the job is deterministic and the user needs to trust the result. Filtering, sorting, validation, calculation, routing with clear rules. Running the cost-and-return test shows exactly when a simpler feature beats an AI feature on ROI, so the comparison is a number rather than a preference. Putting a probabilistic model in front of a job that demands a correct answer is how you manufacture distrust. The model will be confidently wrong some percentage of the time, and that percentage is the whole story for the user.

This is also where the difference from killing a feature matters. If the feature already shipped and is underperforming, that is a separate decision: how to kill an AI feature you already shipped cleanly, after launch, once the data is in. Refusing to build is cheaper than killing, because nothing was spent. The earlier you say no, the smaller the cost of the wrong AI feature is.

Which AI features should I avoid building first

Rank refusals by the cost of being wrong, not by how exciting the idea is. This is a one-line ai cost benefit analysis: the features to avoid building first are the expensive-to-run, low-confidence, trust-dependent ones, because they fail loudest and cost most while failing.

Refusal cost = (run_cost_per_month + maintenance_drag) × months_before_you_admit_it
             − (metric_delta_if_it_worked × confidence_it_will_work)
 
If confidence_it_will_work is a guess, treat it as 0.
Then Refusal cost is pure loss, and the feature ranks first to cut.

The ordering falls out cleanly. Cut the open-ended generative features with high run costs and no named metric before you cut the small supportive nudges that are cheap and reversible. Excitement inverts this ranking, which is exactly why excitement is a bad sorting key.

A refusal applied: scoring one AI idea (Concept Demo)

Take a common request: an AI assistant that writes a customer's project setup for them during onboarding. It demos wonderfully. Run the test.

Metric: activation rate, the share of new accounts reaching first value. Baseline (projected, for illustration): 38% today. The flop-risk scorecard reads metric_named true, baseline_measured true, beats_simpler_tool unknown, trust_required high, adds_a_step false. Trust is the killer here, because users do not trust a setup they did not shape and tend to redo it, which adds time instead of removing it. Projected flop risk lands at 2, in the de-risk band.

So the refusal is not a flat no. It is a sequence: ship the non-AI guided flow with strong defaults first, measure the activation delta, and only then decide whether a generative layer earns a build slot. That is the whole discipline. You did not kill an idea, you refused to build the expensive version of it before the cheap version proved the metric. This is a Concept Demo of the method, not a client result.

Most of the value in deciding which AI features to build is in the refusals, because a roadmap that ships only features tied to a metric and a baseline is a roadmap that compounds instead of bloats. The teams that win the next two years will not be the ones with the longest AI feature list. They will be the ones with the shortest, and the clearest reason for every name on it.

TIP

Want a second opinion on which AI features to build and which to cut before they cost you a sprint? How the AX Audit works.

AI Experience (AX) Audit

Find out which opportunity is actually worth building

The audit looks at your product and your metrics, then tells you where AI earns its place and where it does not.