Most AI features fail before the first line of code

The single biggest risk in adding AI to a product is not picking the wrong model. It is shipping a feature that solves a problem nobody has, on margins that quietly turn negative, wrapped in a UI that users do not trust. The work that matters happens before any code is written.

The numbers are unambiguous. McKinsey's 2024 work found 65% of organisations were regularly using generative AI in at least one function. A PwC survey found only 12.5% of CEOs reported it had delivered both cost savings and revenue growth. By 2025, 42% of companies had scrapped most of their AI initiatives, up from 17% the year before. Gartner expected 30% of generative AI projects to be abandoned after proof-of-concept by the end of 2025. CFOs are asking harder questions now, and "we're experimenting" is no longer a sufficient answer.

Most failures share a root cause: teams start with the technology and work backwards to a problem. There is a widely cited 2025 case of a startup that built a full natural language model for "AI customer support" before testing it with users, then discovered users only needed canned FAQ responses. Eighty per cent of the stack was unnecessary. That pattern repeats everywhere.

A simple test

Write down the problem your AI feature solves without using the word "AI". If the sentence still makes sense and points to a specific user task, you have something. If it collapses into "we want to be more AI-driven", you do not have a feature yet.

The unit economics trap founders walk into

Most product teams are not having this conversation yet. It is the one that will define which AI features survive 2026.

How the margin picture quietly inverts

Take a SaaS product on a flat per-seat plan. You add a "chat with your data" feature using a top-tier model. Most users send small queries. Engagement looks fantastic. Then a subset of accounts start pasting giant CSV exports into the chat window every day. Suddenly that cohort costs more in model and retrieval calls than they pay for the entire product. The headline metrics still look good. The margin curve is bending the wrong way.

This is not hypothetical. AI Dungeon, the text adventure that generated every scene through a language model, became the canonical example of inference costs eating a product alive. When cost scales with usage and the price does not, every user interaction burns cash.

The risk is no longer that you miss the AI boat. The risk is that you board a boat that burns cash with every user interaction.

Mind the Product, 2026 AI Product Strategy

Gating and usage design as a fix, not an afterthought

You can avoid this, but only if you design for it before launch. The questions to answer:

  • What does the top 5% of usage cost you, not the median?
  • Where do you cap, throttle or charge for excess use?
  • Which user actions justify a premium model, and which can route to a cheaper one?
  • Can you cache, batch or pre-compute anything?

These are not engineering decisions made late in the build. They are product decisions made early, and they belong in the same conversation as feature scope. When we work on AI development for clients, this modelling sits at the front of the engagement, not at the end.

Starting as a wrapper is fine. Staying as one is fatal.

A lot of what shipped under the "AI-powered" banner in 2024 and 2025 was a prompt template sitting on top of someone else's API. There is nothing wrong with starting there. The mistake is staying there. If your core feature is one API update away from being eaten by your model provider, you do not have a product. You have a launch announcement.

The defensible position is everything around the model: proprietary data loops, vertical knowledge, integration into workflows users already live in. The 2025 AI Index Report found the capability gap between top models is closing fast, with the Elo score difference between first and tenth place dropping from 11.9% to 5.4% in a year. The provider is not your moat. The integration is.

What Grammarly did that CodeParrot didn't

Grammarly nearly stalled in 2022 with year-over-year revenue growth of 1.44%. They launched GrammarlyGO embedded inside the writing workflows people were already using: Gmail, Docs, Slack. Revenue grew 98.77% the year after, reaching an estimated $700 million ARR by May 2025. The model was not the differentiator. The distribution and workflow depth were.

CodeParrot, a Y Combinator backed tool that converted Figma designs to production code, had a genuinely impressive demo. It shut down by July 2025, outpaced by GitHub Copilot and Replit, who owned the workflow developers already used. A good demo does not beat distribution.

Measure outcomes, not engagement with the AI button

In 2024 it was acceptable to define success as "percentage of users who clicked the AI feature". That was the experimentation phase. It is not enough now.

Useful metrics measure the work the AI is supposed to be doing, not whether users tried it. Compare these:

| Weak metric | Stronger metric | |---|---| | 50% of weekly active users try the assistant | Users who use the assistant open 30% fewer how-to tickets than those who do not | | Adoption rate of AI feature | Time to complete onboarding task X reduced by 40% | | Number of AI sessions per week | Tier 1 support volume on basic questions reduced by 30% |

McKinsey's 2025 work points to workflow redesign as the biggest single driver of EBIT impact from generative AI. The companies extracting value do not have better models. They have better integration into how work actually gets done. This is also why so many enterprise pilots stall. An MIT study found 95% of enterprise gen-AI pilots failed to deliver measurable P&L impact, mostly due to integration, data and governance gaps rather than model capability. McKinsey's 2025 State of AI put 23% of organisations as having scaled AI, with the majority stuck in proof-of-concept.

The fix that works is vertical-first. Pick one end-to-end workflow, deploy AI inside it, prove the ROI completely, then extract reusable components. Building a horizontal AI platform that delivers diffuse value rarely survives a budget review.

Get plain-English guides like this in your inbox.

One short email a month. WordPress, Shopify, SEO, no fluff. Unsubscribe in one click.

We never share your email.

Trust is a design problem, not a model problem

Split comparison of AI interface design: left shows generic output causing user confusion, right shows the same output with confidence levels, source attribution, and control options, resulting in user trust

AI adoption among adults reached 54.6% in 2025, a ten-point jump in twelve months. A 2024 Statista survey found only 46% of consumers were comfortable with brands using AI, down from 57% the year before. People use AI tools personally and distrust brands deploying AI on them. Those two facts coexist without contradiction.

The trust problem is largely a design problem. When something sounds authoritative but is occasionally wrong in confident-sounding language, users feel cognitive dissonance and stop trusting the feature entirely. In 2025, Google's AI Overviews told users to add glue to pizza sauce with the same tone it used for accurate answers. Users had no way to evaluate which outputs were reliable. Trust collapsed for many of them.

Confidence cues, override controls and communicating uncertainty

Models are not equally certain about every output. Most AI features present every output with the same visual weight regardless. That is the design failure. Practical fixes:

  • Show confidence. Surface when the model is less sure, and let users see why.
  • Make reversal easy. Undo, edit and override should be one click, not a hidden setting.
  • Show the working. Where did this answer come from? Which document, which field, which assumption?
  • Fail visibly. When the system does not know, it should say so. Silence or hallucination both kill trust.

Nielsen Norman Group's State of UX 2026 put trust as the dominant design problem for AI experiences this year. Building it requires the same fundamentals as any product: transparency, control, consistency, and recovery when things go wrong. This is unglamorous work. It is also the work that separates features that get adopted from features that get a curious first click and a permanent ignore.

Underneath all of this sits data quality. Air Canada's AI assistant gave a customer incorrect bereavement fare information in February 2024 and was forced to honour the wrong price. The model was not really at fault. The inputs were.

Where the real ROI sits (it's not in the demos)

MIT's work found more than half of generative AI budgets go to sales and marketing tools, but the biggest ROI is coming from back-office automation. Everyone is building where the demos look flashy. The money is where the work is boring. Microsoft reported saving around $500 million in 2025 by introducing AI-powered systems into its call centres.

If you are deciding where to invest, the unglamorous answer is usually right: operations, support, internal tooling, document workflows. For founders building custom software and SaaS, this is where genuine margin improvement lives, even if it makes for a less exciting investor deck.

A practical starting point for your next AI feature

Here is the sequence we use:

  1. Write the problem in one sentence, no "AI" allowed. If you cannot, stop.
  2. Define the outcome metric. Not engagement. The actual work-related number that has to move.
  3. Model the unit economics for the noisiest 5% of users. Not the median.
  4. Decide your gating before launch. Throttling, tiers, premium routing.
  5. Pick the cheapest model that meets the quality bar. Reassess every quarter.
  6. Design for uncertainty. Confidence cues, override paths, transparency about sources.
  7. Pilot vertically. One workflow, end to end, before you generalise.
  8. Measure outcomes, not button clicks.

None of this is about being slow. It is about doing the thinking once, in the right order, so you do not unship the feature in six months. If you want a second pair of eyes on a feature you are scoping now, see how we approach this kind of work. Plenty of teams come to us after a first attempt that demoed well and never quite landed. The second build is usually easier when the constraints are clear from day one.