Glossary - Customers and retention
Churn prediction model
A churn prediction model estimates which customers are likely to stop buying within a set window, giving each one a risk score from their purchase history. In ecommerce, where customers rarely cancel anything and simply go quiet, the model is judged on how many of the customers it flags really do lapse, and on whether acting on its list earns more margin than it costs.
| Variable | Definition |
|---|---|
| Lapsed | No order within the window you define, set from your natural repurchase interval, for example 1.5 to 2 times the typical gap between orders. |
| Base lapse rate | The share of all active customers who lapse in the window with no model at all. |
| Incremental retained | Flagged customers who bought because you acted, measured against a holdout of flagged customers who got nothing. |
The model itself can take several forms. The simplest is a rule: days since last order against the customer's own usual gap. RFM segmentation scores recency, frequency and monetary value together. Probabilistic models such as BG/NBD estimate the chance a customer is still active from their order timing, and machine learning classifiers add signals like discount use, returns and product category. More complex is not automatically better; the right one is the model whose list earns the most margin when you act on it.
Worked example
The model is 2.5 times better than guessing, and the campaign still only just pays. Of the 120 customers who used the discount, 80 would have bought anyway, so two thirds of the discount spend went to people who needed no persuading. The holdout is what reveals that; without it the campaign would have claimed all 120 orders. Precision and recall say whether a model finds the right people. Only the holdout says whether contacting them was worth it, which is the question spotting customers about to churn is really about.
What is a good churn prediction model?
There is no honest accuracy benchmark to borrow, because results depend on how you define lapsed, how regular your repurchase cycle is, and how much history each customer has. A useful model has three traits. It beats the base rate by a clear margin on customers it was not trained on. It flags people early enough that an offer or message can still change the outcome. And it ranks by value at risk, not just probability, so a likely-to-lapse customer with high margin LTV sits above a likely-to-lapse one-time discount buyer.
Accuracy alone is a trap. With a 16% base rate, a model that predicts nobody churns is 84% accurate and completely useless.
Churn prediction vs related metrics
| Metric | What it tells you | How it differs |
|---|---|---|
| Retention rate | The share of customers still buying after a period | Describes the past for a group; a churn model scores each customer's future. |
| RFM segmentation | Customers grouped by recency, frequency and value | A simple, readable churn signal and often the baseline a model must beat. |
| Customer lifecycle | The state each customer is in, from new to lapsed | Churn risk is the probability of moving to the lapsed state next. |
| Win-back | Bringing lapsed customers back | What you do after the model flags someone; judged by incremental margin. |
Common mistakes
- Using a subscription definition of churn. Most ecommerce customers never cancel. Define lapse from each customer's own repurchase rhythm, not a single store-wide cut-off.
- Grading the model on accuracy. With a low base rate, accuracy flatters any model. Use precision, recall and lift, then margin.
- Leaking the future into training. Features built with data from after the prediction date make a model look brilliant in testing and fail in use.
- Acting without a holdout. Without one, every flagged customer who buys looks like a save, and discounts given to loyal buyers look like wins.
- Treating every at-risk customer the same. Ranking by probability alone spends the most effort on customers worth the least. Weight by value at risk, and read whether your VIPs are really profitable.
FAQ
Define lapse from your repurchase interval, take a snapshot date, build features from each customer's history before it (recency, frequency, value, discount use, returns), label who lapsed after it, then train and test on separate customers. Start with an RFM rule as the baseline to beat.
It depends on your data. Probabilistic models like BG/NBD work well from order timing alone. Gradient-boosted classifiers can use richer signals but need more history and care. Pick the one that beats a simple recency rule on held-out customers and earns margin when acted on.
Use precision (how many flagged customers really lapse), recall (how many lapsers were flagged) and lift over the base rate. Plain accuracy misleads when most customers do not lapse. Then run a holdout to measure whether acting on the list adds margin.
Rank them by value at risk, then match the action to the customer: a reminder or new-product message before a discount. Hold back a random share of each group so you can see which actions produce customers who would not have returned anyway.
Updated September 2026