The definition is a decision, not a lookup
There is rarely a column called churned sitting in the database. Somebody has to decide what churn means, and that decision changes the dataset more than any algorithm choice will. The target variable is the one thing in a machine learning project that is entirely yours to define, and it is routinely treated as though it were handed down.
Here is what is at stake. Define churn as “no order in 30 days” and two thirds of your customers are churners. Define it as “no order in 180 days” and three percent are. Same customers, same data, same code, two completely different problems.
Everything downstream inherits this choice: the class balance, which features matter, how far ahead you can predict, and whether the result is actionable. Get it wrong and the modelling work is wasted, no matter how carefully the training loop is executed.
How the window swings your target variable
Watch the positive rate move as the only thing that changes is a number:
import numpy as np
import pandas as pd
rng = np.random.default_rng(9)
customers = 2000
gaps = rng.gamma(shape=1.6, scale=38, size=customers) # days since last order
obs = pd.DataFrame({"customer_id": np.arange(customers),
"days_since_last_order": gaps.round(0)})
for window in [30, 60, 90, 180]:
rate = (obs["days_since_last_order"] > window).mean()
print(window, round(float(rate), 3))
# 30 0.681
# 60 0.399
# 90 0.224
# 180 0.032
A twenty-fold swing in the positive rate, driven by a single integer. At 30 days you are mostly labelling normal purchasing rhythm as churn. At 180 days you have a rare-event problem with 64 positive examples, and a classifier will struggle to find any signal at all.
So which is right? The one that matches the action. If the retention team can only usefully intervene within a month of someone going quiet, a 90-day window labels people who are already gone. If your product is bought twice a year, a 30-day window is noise. Pick the window from the business rhythm and the intervention timing, then check what class balance it leaves you with. If the answer is under about two percent positives, expect to treat it as an imbalanced problem rather than a standard one.
Define the three parts explicitly
Every target definition needs an event, an observation window, and a reference date. “No order placed, within 90 days, measured from the first of the month.” Leave any of the three implicit and two people will build different datasets from the same sentence.
Write it as a filter you could run in SQL. If you cannot, it is not yet a definition.
Censoring: customers who have not had the chance
A new customer who signed up three weeks ago cannot be labelled with a 90-day window. They have not had 90 days. Including them anyway puts rows in your training set whose label is “not churned” purely because time has not passed.
tenure = rng.integers(5, 400, customers) # days since signup
eligible = tenure >= 90
print("eligible for a 90-day label:", int(eligible.sum()), "of", customers)
# eligible for a 90-day label: 1562 of 2000
print(round(float((obs["days_since_last_order"] > 90).mean()), 3))
print(round(float((obs.loc[eligible, "days_since_last_order"] > 90).mean()), 3))
# 0.224
# 0.230
Here the distortion is small, 0.224 against 0.230, because the simulated tenure is roughly uniform. In real data it is often severe: fast-growing companies have a large share of recent signups, and every one of them lands in the negative class by construction. Your model then learns that recently acquired customers do not churn, which is false and will go wrong the moment growth slows.
The fix is unglamorous. Exclude rows whose observation window has not closed. Record how many you dropped and why, because somebody will ask why the row count does not match the customer table.
Temporal structure matters for splitting too, since a random split mixes future rows into training. TimeSeriesSplit is the tool, and Section 7 covers the splitting strategies properly.
Targets that leak, and proxies that mislead
Two failure modes sit specifically in the target, not the features.
The target definition pulls in post-outcome information. If “churn” is derived from a cancellation_reason field that is only populated by the retention team after they have spoken to the customer, then the label encodes the intervention rather than the behaviour. The model will look superb and predict nothing useful, which is the leakage pattern described in the main challenges of machine learning.
The proxy is not the thing you care about. You want to predict customer dissatisfaction; you have complaint tickets. Most dissatisfied customers never file one, and the ones who do are systematically different, so you are modelling complaint-filing behaviour and calling it satisfaction.
Proxies are often unavoidable. The requirement is to name the gap in writing: this target measures complaints raised, which under-represents dissatisfied customers who leave silently. That sentence in your documentation prevents somebody later quoting the model as though it measured the real thing.
A short checklist before you commit
| Check | Question to answer |
|---|---|
| Event | What exactly counts as the outcome occurring? |
| Window | Over what period, measured from when? |
| Eligibility | Which rows have had time to show the outcome? |
| Timing | Is every input to the label known before the prediction moment? |
| Balance | What positive rate does this leave, and is it workable? |
| Proxy gap | What does this measure that is not quite the target? |
Run all six before writing a line of modelling code. The cost of changing your mind at this stage is an afternoon. The cost after you have built features and tuned a model is the whole project. Label timing interacts directly with how supervised learning consumes labelled rows, and the splitting tools are documented under TimeSeriesSplit.
Frequently Asked Questions
How do you define churn for a machine learning model?
Pick an event (no order, no login, subscription cancelled), an observation window that matches your intervention timing, and a reference date. Then exclude customers whose window has not closed. Check the resulting positive rate: under roughly two percent means you are building a rare-event problem.
What is a censored label?
A row where the outcome has not had time to occur. A customer who signed up three weeks ago cannot be labelled against a 90-day churn window. Including them silently adds negatives and teaches the model that recent customers never churn, which fails as soon as growth slows.
What is a proxy target variable?
A measurable stand-in for something you cannot observe directly, such as using complaint tickets to represent dissatisfaction. Proxies are often necessary, but they differ systematically from the real outcome. Document the gap so nobody later reports the model as measuring the underlying thing.
Key Takeaways
- Write your target as an event, a window and a reference date, precise enough that you could express it as a SQL filter.
- Choose the window from how quickly the business can intervene, then check the positive rate it produces rather than accepting whatever falls out.
- Exclude rows whose observation window has not closed, and record the count, or recent records will quietly fill your negative class.
- Trace every field feeding your label and confirm it is populated before the prediction moment, since label leakage produces superb scores and no value.
- Document the gap whenever you use a proxy, so the model is not later quoted as measuring something it never saw.