Which AI features to kill before you ship them

A pre-build kill test for SaaS teams: the signals that say an AI feature should die in planning, before the cost, and which AI features to kill first.

Shahriar P. ShuvoShahriar P. ShuvoAI for SaaS & Features7 min read
Which AI features to kill before you ship them

The cheapest AI feature you will ever build is the one you kill in planning. It costs a meeting. The most expensive one is the feature you ship, watch underperform for two quarters, and quietly remove after it has already taken engineering time, model spend, and a slot on your roadmap.

Most teams get this backwards. They audit AI features after launch, once the numbers are in. The real leverage is upstream, in the list of ideas you decide not to build at all. So this piece is about the AI features to kill, and the signals that tell you to kill them before a line of code exists. You will leave with a kill test you can run on your backlog this week.

This is the inverse of the usual question. Instead of asking which AI features earn their place, defined in our cornerstone on which AI features earn their place, we ask which ones never should have been on the list.

Why most AI features should die in planning, not production

A feature should die in planning when it has no metric, no demand, and no path to user trust. All three are knowable before you build. None of them require a prototype to assess.

The failure data backs this up. Gartner projects that at least 30% of generative AI projects will be abandoned after proof of concept by the end of 2025, citing unclear business value and escalating costs among the top reasons. "Unclear business value" is a planning failure wearing a production costume. The value was never defined, so the project drifted until someone pulled the plug.

At least 30% of generative AI projects will be abandoned after proof of concept by the end of 2025, due to poor data quality, inadequate risk controls, escalating costs or unclear business value. (Gartner, 2024)

RAND studied why AI projects fail and found five leading root causes: teams misunderstand the problem, lack the data to train a model, chase the newest technology over a real user need, run on inadequate infrastructure, or aim at a problem too hard for AI. Four of those five are visible in a planning review. You do not need to ship to know your data is thin or your problem is fuzzy. You need to be honest in a room.

That honesty is the whole game. The kill is the cheapest decision on the table, and it is the one teams avoid because saying no to an AI idea feels like saying no to the future. It is not. It is saying no to spending six figures to learn something a one-hour review could have told you.

What are signs an AI feature will flop

There are five signals. Any one of them is a yellow flag. Two or more is a kill. Run every idea in your backlog through this table.

SignalThe question to askKill threshold
No metricWhich number that we already track does this move?No named metric, or "engagement" with no definition
No demandHave users asked for this, or are we guessing?Zero pull signal; the idea came from a demo, not a user
No trust pathWhat happens when the model is confidently wrong?No fallback, no human-in-the-loop, no way to recover
Underwater economicsWhat does each successful use cost in tokens and review?Per-use cost exceeds the value of the action it replaces
A simpler feature winsCould a rule, a filter, or a form do 80% of this?A non-AI version ships faster and moves the metric more

The metric signal is the one most teams skip. If you cannot name the number a feature is supposed to move, before you build it, you have no way to tell success from noise after you ship. That is the same discipline behind setting a baseline first: a feature with no baseline is unkillable, because you can never prove it failed.

The economics signal kills more good-looking ideas than any other. A feature can be genuinely useful and still lose money on every call. We go deeper on the cost of the wrong feature, but the short version is that token cost plus human review per successful action is your real unit cost, and it is rarely on the slide.

Which AI features should you not build

Some patterns flop reliably enough to name. These are the AI product features you should not build, regardless of how good the demo looks.

  • The demo-driven feature. It exists because it was fun to prototype, not because a user asked. It has no owner and no metric. This is feature creep with a model attached: complexity added, value assumed.
  • The me-too chatbot. A generic assistant bolted onto a product that did not need one. It competes for attention with the workflows people actually came for, and it is the category most prone to escalating cost. Gartner expects over 40% of agentic AI projects to be canceled by the end of 2027, citing the same escalating costs and unclear value.
  • The feature that fights your core engine. If your product already does the job deterministically, wrapping it in a probabilistic model usually adds risk, not value. Supportive AI sits above a reliable core; it does not replace it.
  • The orphan. No team owns the metric, so no one is accountable for whether it works. Orphan features survive because killing them is awkward, not because they earn their place.

Lists of the best AI features for SaaS will tell you what to add. They rarely tell you what to cut. Our companion piece on the features you should not build is the other half of that map, and the two together are how you keep a backlog honest.

WARNING

Define the metric before the model. An AI feature with no number attached is not a feature, it is a bet you cannot grade. If you would not fund it as a non-AI feature, the model does not change the math.

How do you decide to cut an AI feature

You decide by running the idea through a gate, not by debating it. The gate is mechanical on purpose, so the kill is a number and not an argument about vision. Score each idea against five gates. If two or more fail, it dies in planning.

KILL TEST  (run before any build)
 
gate 1  metric:     names a number we already track?        [pass/fail]
gate 2  baseline:   we know its current value today?         [pass/fail]
gate 3  delta:      projected movement is worth the cost?    [pass/fail]
gate 4  economics:  per-use cost < value of the action?      [pass/fail]
gate 5  trust:      a safe path when the model is wrong?      [pass/fail]
 
rule:   2+ fails  -> kill in planning
        1 fail    -> fix the gap or descope, then re-score
        0 fails   -> build the smallest version, measure the delta

Gate 3 is where projection does the work. You do not need real numbers to run this; you need an honest estimate of the delta against the metric, framed as projected, then weighed against projected cost. This is the same scoring logic behind deciding which AI features to build, just pointed at removal instead of addition. A feature that cannot clear a projected go decision will not clear a real one.

The reason to make it mechanical is that AI feature adoption is emotional. Teams fall in love with demos. A gate does not. It asks the same five questions of every idea, and it lets you kill the founder's pet feature with the same straight face you use on everyone else's.

When killing the AI feature is the wrong call

Kill the feature, not the bet. Sometimes a feature looks dead because the metric was wrong, the baseline was never set, or the model had no trust path, not because the idea had no merit. Those are fixable gaps, and fixing one can revive a feature that the kill test flagged.

A feature that fails only the trust gate is often a reliability problem, not a value problem. Add a confidence threshold, a human-in-the-loop step, or a graceful fallback, and the same feature can clear the gate on a re-score. Weak AI feature adoption frequently traces to a missing trust path rather than a missing use case.

The clean way to handle the ambiguous cases is to ship the smallest possible version, set the baseline, and let the metric decide. That is a different discipline from the planning kill, and we cover it in killing an AI feature after launch. The two are complementary: the planning kill saves you from building the obvious flops, and the post-launch kill saves you from keeping the ones that looked fine on paper and moved nothing.

The AI features to kill, applied: a worked example (Concept Demo)

Here is the gate applied to two ideas, framed as projected and designed to move. No real client, no claimed result.

IdeaMetricBaselineProjected delta vs costTrust pathVerdict
AI churn-risk summary in the account viewLogo churnKnown (tracked monthly)Worth it: cheap per-use, ties to a tracked numberRead-only, human acts on itBuild small, measure
Generic in-app chat assistant"Engagement" (undefined)UnknownUnderwater: high token cost, fuzzy valueConfidently wrong, no fallbackKill in planning

The first idea passes because it points at a number you already watch and degrades safely. The second fails three gates before anyone writes code. That is the kill test earning its place: it costs an hour and saves a quarter.

Build the discipline once and it compounds. Every clean kill frees a roadmap slot for a feature that moves a number, and a shorter list of the right AI features to kill is what keeps your roadmap pointed at value instead of theater. The teams that ship AI well are not the ones that build the most; they are the ones that say no the earliest.

TIP

Not sure which AI features to kill and which earn their place? How the AX Audit works. We run your backlog through the gate and project the delta on a metric you already track.

AI Redesign & Rescue

Your AI feature is live. Nobody uses it.

We rebuild the part worth keeping and remove the part that was never going to work, inside the product you already shipped.