Survival Analysis for Churn: Beyond Binary Classification

Codeayan Team · Oct 7, 2026 · 0 Views

Take 2,000 subscribers on ₹299, ₹499 and ₹999 plans and watch them for twelve months. 685 cancel. The other 1,315 are still paying when the window closes. A binary churn classifier has to file those 1,315 under “did not churn”, which is wrong for everyone who cancels in month thirteen. Survival analysis exists for exactly this gap.

Short answer

Survival analysis is a family of statistical methods that model the time until an event, such as a customer cancelling, while treating customers who have not yet cancelled as censored rather than as non-churners. A churn classifier answers whether a customer leaves inside a fixed window. A survival model answers when, for every customer, at every month.

Why does a binary churn classifier lose information?

Because the label is a snapshot. A model trained on “churned within 12 months” scores a customer who cancels in week two the same as one who cancels in month eleven, and it files every active customer under “no churn”, even the ones leaving next week. Timing and censoring both vanish.

The fix is to keep two numbers per customer: how long you observed them, and whether the event happened. Anyone still active at the cutoff is right-censored. You know they lasted at least that long, nothing more. Censoring is the norm. In the Rossi recidivism dataset bundled with lifelines, 318 of 432 records are right-censored. The standard tools are the Kaplan-Meier estimator (1958) for a population curve and the Cox proportional hazards model (1972) for covariates. A logistic regression baseline still has a place, but it answers a narrower question.

PropertyBinary classifierSurvival model
Target 1 if churned within N months. Time observed, plus an event flag.
Active customers Forced to class 0, or dropped. Censored; they still contribute exposure time.
Output One probability for one window. A survival curve S(t) for any month.
Main metric AUC at the chosen window. Concordance index, plus calibration at chosen times.

How do you run survival analysis for churn in Python?

Survival analysis in Python takes two fits, Kaplan-Meier first and then a Cox model, both in lifelines. Each row needs a duration column and an event flag, and the library handles the censoring bookkeeping. The snippet simulates the 2,000-subscriber cohort from above.

Python: Kaplan-Meier and Cox on a Simulated Cohort
1import numpy as np
2import pandas as pd
3from lifelines import KaplanMeierFitter, CoxPHFitter
4
5# simulated cohort: 12-month window, active customers are censored
6rng = np.random.default_rng(7)
7n = 2000
8plan_inr = rng.choice([299, 499, 999], n)
9support_tickets = rng.poisson(2, n)
10rate = 0.03 * np.exp(0.002 * (plan_inr - 499) + 0.15 * support_tickets - 0.5)
11tenure = rng.exponential(1 / rate)
12df = pd.DataFrame({
13    "months": np.minimum(tenure, 12).round(2),
14    "churned": (tenure <= 12).astype(int),
15    "plan_inr": plan_inr,
16    "support_tickets": support_tickets,
17})
18
19kmf = KaplanMeierFitter().fit(df.months, df.churned)
20print(kmf.predict(6))  # share still subscribed at month 6
21
22cph = CoxPHFitter().fit(df, duration_col="months", event_col="churned")
23cph.check_assumptions(df, p_value_threshold=0.05, show_plots=False)
24print(cph.predict_survival_function(df.head(3), times=[6]))

In this simulated cohort, 66% of rows are censored and Kaplan-Meier puts month-6 retention near 81%. The Cox fit returns a hazard ratio of about 1.19 for support_tickets: each extra ticket multiplies a customer’s monthly churn hazard by roughly 1.19. That is a rate ratio, not an odds ratio, so coefficient interpretation needs the survival framing. The last line returns S(6) per customer. Rank by 1 minus that value, or pick month 3 or month 9, without retraining anything.

What breaks a survival model in production?

Three things, in order of how often they bite: the proportional hazards assumption, covariates that change over time, and a concordance score that hides poor calibration. The Cox model assumes each covariate’s effect on the hazard stays constant across tenure. Test it with check_assumptions(). If it fails, stratify with strata= or switch to WeibullAFTFitter.

Support tickets are the textbook violator. A customer’s count at month 10 is not their count at signup, and a static column quietly leaks the future. Use CoxTimeVaryingFitter with long-format rows instead. Then evaluation. The concordance index only ranks customers; it never says whether “30% chance of churning by month 6” is true. And the scikit-survival documentation on evaluating survival models describes a bias in Harrell’s concordance index under heavy censoring, with Uno’s estimator as the alternative. Check calibration at the horizon your retention team acts on, using the classification metrics habit, adjusted for censoring.

Do not drop the censored rows. Labelling still-active customers 0 understates churn, and discarding them overstates it, because only the early leavers remain. The first error looks like good news in a dashboard, which is why it survives review.

  • Survival analysis models time to an event and keeps censored customers as partial information instead of mislabelling them.
  • A binary churn classifier returns one probability for one fixed window; a survival model returns a curve S(t) for every customer.
  • The Kaplan-Meier estimator gives a population retention curve, and the Cox proportional hazards model adds covariates through hazard ratios.
  • The Cox model assumes constant covariate effects over tenure, so run check_assumptions() before trusting any hazard ratio.
  • Labelling active customers as non-churners understates churn; dropping them overstates it.

Conclusion

Run Kaplan-Meier on your last cohort and compare month-6 retention with what your classifier implied. The gap is the price of ignoring censoring. Once the Cox model is fitted, interpreting its coefficients is the next skill to build.

Frequently Asked Questions

What is survival analysis in churn prediction?

Survival analysis models how long customers stay before cancelling, instead of only whether they cancel inside a fixed window. Each customer contributes a duration and an event flag. Customers still active at the cutoff count as censored, so the model uses their observed tenure without assuming they will never leave.

What is censoring in customer churn data?

Censoring means you know a customer stayed at least until your last observation, but not when, or whether, they will leave. Every active subscriber at the data cutoff is right-censored. Survival models use that partial information correctly, while a standard classifier must mislabel those customers or discard them.

Why not just use logistic regression for churn?

Logistic regression predicts one probability for one fixed window and ignores when churn happens inside it. It also forces active customers into class 0 or drops them. It remains a fair baseline for a single decision horizon, but it cannot produce a retention curve for arbitrary months.

What does a hazard ratio of 1.2 mean in a Cox model?

A hazard ratio of 1.2 means a one-unit increase in that covariate multiplies the instantaneous churn rate by 1.2, a 20% higher hazard at any tenure, holding other covariates fixed. It is a rate ratio, not an odds ratio, and it assumes the effect stays constant over time.

Which Python library should I use for survival analysis?

Use lifelines for interpretable work: it offers Kaplan-Meier, Cox regression, Weibull accelerated failure time models, and assumption checks through a pandas-friendly API. Use scikit-survival when you need scikit-learn pipelines, penalized Cox models, random survival forests, or censoring-aware metrics such as Uno’s concordance index.

What is the proportional hazards assumption?

The proportional hazards assumption says a covariate multiplies the baseline churn hazard by the same factor at every point in a customer’s tenure. If a feature matters early but fades later, the assumption fails. lifelines flags violations with check_assumptions(), and stratification or an accelerated failure time model are the usual fixes.