AI Adopted. Retention Unmoved. The Metric Nobody Tracks.
TL;DR
- McKinsey's 2025 survey: 88% of organizations use AI, but only 39% can point to profit from it. The gap is widening, not closing.
- Subscription brands are the worst offenders, because retention is the whole business - and they're measuring the AI, not the revenue it retains.
- An "85% accurate" churn model sounds like a win. The number that matters is how many subscribers it actually keeps paying. Usually a fraction of what the dashboard claims.
- Four rules to close the gap - and permission to kill the AI that doesn't move cohort revenue.
Two numbers that don't fit together 🤖
McKinsey's 2025 AI survey landed on a pair of numbers that should make any operator uncomfortable. 88% of organizations now use AI in at least one function, up from 78% a year earlier. Only 39% report any measurable impact on their bottom line.
That gap - nearly 50 points between adoption and proven profit - isn't shrinking. Gartner's reading the same pattern: fewer than one in ten AI investments ever reach full production. Most stall in pilot purgatory, generating dashboards and slide decks instead of money.
For subscription brands this stings more than most, because retention is the business. Acquisition is a one-time transaction. Retention is every transaction after that, compounding. So you'd think the category that lives and dies by retention would be the one nailing AI-driven retention. Mostly, it isn't.
What "using AI" actually means on your team 📉
Walk through what a typical subscription team has bolted on over the last two years. There's a churn-scoring model flagging at-risk subscribers. A recommendation widget suggests the next product. Your chatbot deflects cancellation requests. A win-back flow fires the moment someone pauses.
Each one ships with a dashboard. The churn model reports accuracy - 85%, 92%, pick a number. The recommendation widget reports click-through. Your chatbot tracks deflection rate. The win-back flow logs open rate.
Every one of those is an activity metric. It tells you the tool is doing something. None of them tell you whether more subscribers are still paying you next month than would have been otherwise. You bought a retention tool and you're measuring the tool.
The math that exposes the gap 📈
Here's where the "85% accurate" story falls apart. Run it with me, then swap in your own numbers.
Say you have 50,000 active subscribers and monthly churn of 5%. You lose 2,500 subscribers a month. You deploy an AI churn model. The vendor's headline: 85% recall - it flags 85% of would-be churners before they cancel. Sounds like you've solved half the problem.
Flagging a subscriber doesn't save them. You still have to reach them, and the reach has to work. Say the model flags roughly 2,130 would-be churners. You send each one a win-back offer. A generous save rate on a win-back touch is around 15%. That keeps about 319 subscribers who would have left.
319 out of 2,500. Your churn drops from 5.0% to roughly 4.9%. A tenth of a point.
If your average subscriber pays $25 a month, that's around $8,000 retained monthly, roughly $96,000 annualized. That can absolutely be worth it - depending on what the tool costs and what the offers cost you in margin. The math can be fine. What you report to your board is "85% model accuracy," and what actually happened is 319 subscribers. The dashboard makes the gap invisible.
The metric that's lying to you 🎯
Model accuracy is the most seductive number in retention tech, because it's big and it sounds scientific. It's also nearly useless on its own. A model can be 95% accurate at flagging who will churn and still save you almost nobody, if your intervention is weak or your offers are generic.
The number that tells you whether your AI retention stack is working is cohort revenue. Take the subscribers the model flagged and you intervened on. Compare their 30-, 60-, 90-day retained revenue to a matched cohort you didn't touch. That delta is the entire justification for the tool. Everything else is a proxy.
Most teams don't run that comparison because it's harder than reading a dashboard and the answer might be uncomfortable. The AI can be "working" by its own metrics and flat on retained revenue at the same time. Until you measure the second thing, you're spending on faith.
The fix: close the gap 💡
Four rules. None of them require new tooling - they require a different question.
- Tie every AI intervention to a cohort revenue metric. Before you switch on a churn model, a recommendation engine, or a win-back flow, define the cohort and the revenue delta you're trying to move. If you can't state the metric in one sentence, you're not ready to measure it. "Increase 90-day retained revenue on flagged subscribers by X" is a metric. "Improve model accuracy" is a vendor's metric.
- Keep a holdout group. The cleanest way to know if your AI is earning its keep is to not run it on everyone. Hold back a matched slice of subscribers - 10%, 20% - and compare their retention to the treated group. Without a holdout, every improvement looks like the AI's work, including the ones caused by seasonality, a price change, or pure luck.
- Price the full intervention, not just the tool. The subscription line item is the easy part. The real cost is the offers you send, the discount margin you give up, the engineering hours to integrate, and the human hours to manage the flows. A tool that "retains" $96,000 a year isn't a win if the offers and ops cost $110,000 to deliver. The unit economics include the save.
- Kill the AI that doesn't move the number. This is the rule nobody follows. If a cohort comparison shows your chatbot or your recommendation widget isn't lifting retained revenue after 90 days, turn it off. Pilot purgatory exists because killing a pilot feels like failure. Keeping a neutral tool running is the actual failure - it's spent budget and attention that could go somewhere with a real delta.
The honest caveat
Not every AI investment needs to prove revenue on a 90-day clock. Some of it is defensive - fraud detection, payment routing, support deflection that shows up as cost savings rather than retained revenue. Some of it is foundational - data infrastructure that makes the next three interventions possible. Those earn a different kind of patience.
But the core claim is uncomfortable and worth sitting with: if you've deployed AI across your retention stack and your cohort retention curve looks the same as it did a year ago, the AI is busy and your business isn't better off. The McKinsey gap shows up right here, between the dashboard you check and the revenue line you own.
After 11 years running retention at Scentbird, the teams I saw win with technology were never the ones with the most tools. They were the ones who could point at a cohort and say, with numbers, "this is what changed when we turned it on." (That discipline is most of what we built into Finsi - making the cohort delta legible enough that you stop trusting the dashboard and start trusting the revenue.)
So here's the question worth taking seriously: if you switched off every AI tool in your retention stack tomorrow, which one would your revenue actually miss?