An objective is not a prediction task
“Reduce churn” is an objective. It is not something a model can output. Machine learning problem framing is the translation step between the two, and it is the stage where most of a project’s eventual success or failure is decided, well before any estimator is chosen.
The translation has a specific shape. You need a unit of prediction, a task type, a decision the output will change, and a number the model has to beat. Get those four written down and the modelling becomes routine. Skip them and you will produce something accurate and unusable.
Business framing, use cases and stakeholder questions belong to the end-to-end predictive modelling workflow. This chapter takes the engineering half: how the objective becomes an X and a y with a threshold attached.
Four questions machine learning problem framing has to answer
What is one row?
The unit of prediction determines the shape of everything downstream. “Reduce churn” could mean one row per customer, one row per customer per month, or one row per subscription. These are different datasets with different row counts, different leakage risks and different targets.
Pick the unit that matches the decision. If the retention team runs a campaign monthly, the unit is customer-month, because that is when the decision is made. Choosing customer-level when the decision repeats monthly throws away most of your training data and gives you a model that cannot say when.
What type of prediction?
The answer is a category, a number, or an ordering. Those map to classification, regression and ranking, and the choice follows from the label’s type, as covered in supervised learning.
Watch for problems that look like classification and are really ranking. “Which customers will churn” invites a yes/no model. If the actual use is “call the 200 most at-risk customers this week”, you need an ordering, and the cutoff is set by call-centre capacity rather than by a probability of 0.5.
What decision changes?
Write the sentence: when the model outputs X, somebody does Y instead of Z. If you cannot complete it, stop. A prediction that changes nothing is an expensive dashboard.
This question also fixes your latency budget and your feature set. A decision made at checkout needs a prediction in under 200 milliseconds using only data available at that instant. A decision made in a Monday planning meeting can take an hour and use everything.
What does being wrong cost?
False positives and false negatives almost never cost the same. Contacting a happy customer with a retention offer costs a discount. Missing a leaver costs their remaining lifetime value. When the ratio is ten to one, a model optimised for balanced accuracy is optimising the wrong thing.
The baseline the model has to beat
Before fitting anything, establish what doing nothing scores. Without it you cannot tell whether your model is good.
from sklearn.dummy import DummyClassifier
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score, roc_auc_score
# loans, X and y as built in the previous chapter
Xtr, Xte, ytr, yte = train_test_split(
X, y, test_size=0.3, random_state=4, stratify=y)
dummy = DummyClassifier(strategy="most_frequent").fit(Xtr, ytr)
print("dummy accuracy", round(accuracy_score(yte, dummy.predict(Xte)), 3))
# dummy accuracy 0.822
pipe.fit(Xtr, ytr)
print("model accuracy", round(accuracy_score(yte, pipe.predict(Xte)), 3))
print("model AUC", round(roc_auc_score(yte, pipe.predict_proba(Xte)[:, 1]), 3))
# model accuracy 0.885
# model AUC 0.926
A model that predicts “never defaults” for every applicant scores 82.2%. The real model scores 88.5%. Reported alone, 88.5% sounds strong; reported next to 82.2% it is six points of actual value, and whether six points justifies the project is a business question you can now ask properly.
Three baselines are worth computing, and the strongest is usually the third: the majority class, a single best feature thresholded by hand, and whatever rule the business currently uses. Beating a DummyClassifier is easy. Beating the senior underwriter’s rule of thumb is the real test.
Capacity changes the question entirely
Most predictions feed an action with a limited budget, and that changes what you should optimise.
import numpy as np
scores = pipe.predict_proba(Xte)[:, 1]
order = np.argsort(scores)[::-1] # highest risk first
actual = yte.to_numpy()
for capacity in [20, 50, 100]:
caught = int(actual[order[:capacity]].sum())
print(capacity, caught, round(caught / actual.sum(), 3))
# 20 15 0.312
# 50 33 0.688
# 100 46 0.958
print("total defaulters", int(actual.sum()), "of", len(actual))
# total defaulters 48 of 270
Reviewing the 50 riskiest applications out of 270 catches 33 of the 48 eventual defaults. Reviewing 100 catches 46 of 48. The second half of that review capacity buys 13 more catches for 50 more reviews, and whether that trade is worth making is the actual decision.
Notice that no threshold of 0.5 appears anywhere. The question was never “is this applicant a defaulter”. It was “given that we can review 50 files this week, which 50”. Framing it that way changes the metric you report, and often the model you pick.
When the answer is not a model
Say so early, because it is cheaper than discovering it in week six. The honest cases: the rule is already written down somewhere and only needs implementing; the decision is made once rather than thousands of times; nobody can act on the output within the window where it is valid; or no labelled history exists and creating it would cost more than the problem is worth. The system taxonomy helps here, since a problem with no labels and no simulator rules out most of the field immediately.
One more case, less often admitted. Sometimes the model is feasible and the organisation cannot absorb it, because nobody owns the resulting action. Framing should surface that, and the test is the sentence from earlier: when the model outputs X, somebody does Y. If no named team does Y, the project is not ready, and no amount of lifecycle discipline will rescue it. Baseline strategies are documented under DummyClassifier.
Frequently Asked Questions
How do you turn a business problem into a machine learning problem?
Specify four things: what one row represents, whether the output is a category, a number or an ordering, which decision the output changes, and what the two kinds of error cost. Then set the number your model must beat. If any of the four is unanswerable, the problem is not ready.
What is a baseline model and why do you need one?
A baseline is the simplest thing that makes predictions, usually the majority class or the rule currently in use. It tells you how much of your model’s score comes from the model rather than the data’s natural balance. An 88% accuracy means little against an 82% baseline.
Should I use classification or ranking for churn prediction?
Ranking, in most cases. If the retention team can contact a fixed number of customers, you need an ordering by risk and a cutoff set by capacity, not a yes/no answer at probability 0.5. Classification suits cases where every predicted positive triggers an action automatically.
Key Takeaways
- Choose the unit of prediction to match when the decision is made, since a customer-level row cannot answer a question the business asks monthly.
- Write the sentence “when the model outputs X, somebody does Y instead of Z” before modelling, and stop the project if you cannot complete it.
- Compute the current business rule as a baseline, not just a majority-class dummy, because that is the comparison anyone will actually challenge you on.
- Set your decision threshold from the capacity of the team acting on the output rather than from a default probability of 0.5.
- State the cost ratio between false positives and false negatives during framing, as it determines which metric is honest for this problem.