Data science · 10 min read
LTV forecasting with under a year of data
Which customer lifetime value models hold up on short purchase histories, how to sanity-check them, and when a cohort table beats all of them.
Every growth conversation eventually asks the same question: how much can we pay to acquire a customer? The honest answer depends on lifetime value, and young businesses rarely have the history to measure it directly. Here is how we forecast it anyway, without fooling ourselves.
Why short histories break naive LTV
The common shortcut — average order value × purchase frequency × an assumed lifetime — hides three problems. Frequency is measured over a window shorter than a customer’s real buying cycle. “Lifetime” is a guess. And the average blends one-time buyers with loyal repeaters, whose behaviour is completely different.
The result is a single number that looks precise and is usually wrong in a predictable direction: too low for stores with slow repeat cycles, too high right after a promotion brought in bargain hunters.
The model we start with
For non-subscription businesses we start with two probabilistic models that work together:
- BG/NBD (or Pareto/NBD) models how many purchases each customer will make. It assumes each customer buys at their own rate while “alive”, and may silently churn after any purchase.
- Gamma-Gamma models how much each customer spends per order, shrinking customers with few orders towards the population average.
Their output multiplied together — and optionally discounted — gives an expected customer value over a chosen horizon. Both need only four columns per customer:
| Field | Meaning |
|---|---|
| frequency | number of repeat purchases |
| recency | time between first and last purchase |
| T | time since first purchase |
| monetary value | average value of repeat orders |
In Python, the lifetimes library or PyMC-Marketing’s CLV module fit both models in a few lines.
Validate before you believe it
The step most teams skip: split the history in two. Fit the model on a calibration period, predict purchases in the holdout period, and compare with what customers actually did.
We look at three checks:
- Total repeat purchases in the holdout: predicted against actual, week by week.
- Calibration against holdout by frequency bucket: customers with 0, 1, 2… calibration repeats should line up with their actual holdout purchases.
- Cohort sanity check: per acquisition month, does the forecast revenue curve continue the observed cohort curve smoothly?
If the model fails these, it is usually telling you something about the data — a promotion that pulled purchases forward, a returns problem, wholesale orders mixed in with retail — rather than about the model.
Keep the horizon honest
A model fitted on eight months of history can say something useful about the next twelve to sixteen. It cannot say much about year three. We cap the reported horizon at roughly twice the observed history, show intervals rather than point estimates, and re-fit monthly as data accumulates.
When a cohort table is enough
If most revenue comes from first orders and repeat rates are low, a plain cohort revenue table — cumulative revenue per customer by months since acquisition — answers the budget question just as well, and everyone in the room understands it. The probabilistic model earns its keep when repeat purchasing matters and you need customer-level predictions: for segmenting, for retention campaigns, or for bidding on high-value lookalikes.
Turning the forecast into decisions
The forecast is only useful once it is attached to spend. We report predicted 12-month value per acquisition cohort and channel next to the cost of acquiring that cohort, and flag channels whose payback period exceeds what the business can finance. That table, refreshed monthly, is usually the most-read output of the whole project.
Key takeaways
- Split your history into calibration and holdout periods; a model that hasn't been checked on a holdout isn't a forecast.
- BG/NBD plus Gamma-Gamma is a strong default for non-subscription stores with repeat purchases.
- Keep the forecast horizon short — about twice the observed history — and widen intervals beyond that.
- Always compare the model with a plain cohort revenue curve; large disagreement means a data or assumption problem.
FAQ
What is the best LTV model for a young e-commerce store?
For non-subscription stores, a BG/NBD (or Pareto/NBD) model for the number of future purchases combined with a Gamma-Gamma model for average order value is a reliable default. It only needs each customer's purchase frequency, recency, age and average order value.
How much data do you need to forecast customer lifetime value?
A few months of transactions with a meaningful share of repeat buyers is enough to fit a probabilistic model, but forecasts should be limited to roughly twice the length of the observed history and validated on a holdout period.
Should I use machine learning instead of BG/NBD for LTV?
Gradient-boosted models can beat probabilistic models when you have long histories and rich features such as browsing, product and channel data. With under a year of data they tend to overfit; start with BG/NBD and Gamma-Gamma as a baseline.