The bug that makes your model look brilliant
Data leakage is information reaching the model during training that will not be available when it predicts. It is the most expensive mistake in applied machine learning, and the reason is structural: every other bug makes your scores worse, so you investigate. This one makes them better, so you celebrate and ship.
Two numbers from this chapter. A single post-outcome column lifted a fraud model from 0.909 to 0.986 AUC. And on a dataset of pure random noise, with no relationship to the target whatsoever, selecting features before cross-validation produced an AUC of 0.812 instead of the 0.5 that honest evaluation gives.
Both look like success. Neither model can predict anything.
Four kinds of data leakage, and how each gets in
Target leakage. A feature that is a consequence of the outcome rather than a cause. The classic shape is a column populated by a process that only runs after the event.
import numpy as np
import pandas as pd
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.model_selection import cross_val_score
# eng, features and y from chapter 4.6
leaky = eng.copy()
leaky["fraud_case_opened"] = np.where(
leaky["is_fraud"] == 1, rng.random(m) < 0.92, rng.random(m) < 0.03).astype(int)
def auc(frame, cols):
pipe = make_pipeline(StandardScaler(), LogisticRegression(max_iter=3000))
return round(cross_val_score(pipe, frame[cols], y, cv=5, scoring="roc_auc").mean(), 4)
print("with fraud_case_opened:", auc(leaky, features + ["fraud_case_opened"]))
print("without it: ", auc(leaky, features))
# with fraud_case_opened: 0.9855
# without it: 0.9089
The investigations team opens a case because fraud was detected. Including that column means predicting fraud from the response to fraud, which is unavailable at the moment a transaction needs scoring.
Train-test contamination. Any transformation that learns from data and is applied before splitting. Imputing with a mean computed on everything, scaling with a standard deviation computed on everything, or selecting columns by looking at the whole dataset.
Temporal leakage. Training on rows that occurred after the rows you test on. A random split on time-ordered data guarantees this, and the fix is a point-in-time construction of the kind described in collecting and sourcing data.
Group leakage. The same entity appearing on both sides of a split. Three transactions from one customer, two in training and one in test, means the model has partly seen the answer. Use GroupKFold when rows cluster by customer, patient or device.
| Kind | Typical cause | Fix |
|---|---|---|
| Target leakage | Column populated after the outcome | Audit when each field is written |
| Contamination | Preprocessing fitted before the split | Put every step in a Pipeline |
| Temporal | Random split on time-ordered rows | Split by date, use TimeSeriesSplit |
| Group | Same entity in train and test | GroupKFold on the entity id |
| Duplicates | Repeated rows across the split | Deduplicate before splitting |
How large the contamination actually is
This is worth seeing rather than being told about, because the magnitude is hard to believe. The data below contains no signal at all: 300 rows, 2000 random columns, and a target assigned by coin flip.
from sklearn.feature_selection import SelectKBest
rng2 = np.random.default_rng(7)
X_noise = pd.DataFrame(rng2.normal(0, 1, (300, 2000)))
y_noise = pd.Series(rng2.integers(0, 2, 300)) # pure coin flips
# the wrong way: pick the 20 best columns using all the data, then evaluate
corr = X_noise.apply(lambda c: abs(np.corrcoef(c, y_noise)[0, 1]))
best = corr.sort_values(ascending=False).head(20).index
print(round(cross_val_score(LogisticRegression(max_iter=3000),
X_noise[best], y_noise, cv=5, scoring="roc_auc").mean(), 4))
# 0.8119
# the right way: select inside the pipeline, refitted on each fold
def corr_score(Xf, yf):
return np.abs([np.corrcoef(Xf[:, j], yf)[0, 1] for j in range(Xf.shape[1])])
inside = make_pipeline(SelectKBest(corr_score, k=20),
LogisticRegression(max_iter=3000))
print(round(cross_val_score(inside, X_noise.to_numpy(), y_noise.to_numpy(),
cv=5, scoring="roc_auc").mean(), 4))
# 0.4242
An AUC of 0.812 on coin flips. With 2000 random columns and 300 rows, twenty of them will correlate with the target by chance, and choosing them using the full dataset means every fold is tested on rows that helped pick those columns. Done correctly, the same procedure scores 0.424, which is the honest answer for data containing nothing.
This is why the selection advice in feature selection techniques insisted on putting selectors inside the pipeline. The gap here is 0.39 AUC, manufactured entirely by the order of two operations.
Detecting it before production does
Leakage has a signature, and you can look for it deliberately.
The score is too good. A model at 0.99 on a problem humans find hard has found a shortcut. Treat any result that would make the project a success on the first attempt as a reason to audit, not to present.
One feature dominates. When a single column holds most of the importance and removing it collapses the model, ask when that column is written. The importance tooling from reading a model’s output is how you spot it.
Validation and production disagree. Offline AUC of 0.95, live performance near random. By this point the cost has been paid.
The audit that prevents most of it is one question per column: at the moment this prediction must be made, has this value already been written? Not “could it be computed”, but “was it, in the source system, before the outcome”. Columns derived from case notes, resolution codes, cancellation reasons and anything a human entered after reviewing the case all fail this test.
A short checklist to apply before every modelling run. Deduplicate, then split. Split by time when rows are time-ordered, and by group when rows cluster by entity. Put every imputer, scaler, encoder and selector inside a Pipeline. Audit each column’s write time against the prediction moment. Then treat a suspiciously high score as a finding to investigate rather than a result, which is the stance recommended in the main challenges of machine learning. scikit-learn’s guide to common pitfalls and data leakage documents the preprocessing cases in detail.
Frequently Asked Questions
What is data leakage in machine learning?
Information available during training that will not exist when the model predicts. It comes from features recorded after the outcome, preprocessing fitted before the train test split, random splits on time-ordered data, or the same entity appearing in both halves. The result is high scores that collapse in production.
How do you detect data leakage?
Look for scores that are too good for the problem, a single feature dominating the model, or a large gap between validation and live performance. Then audit each column by asking whether its value was written before the prediction moment, not merely whether it could be computed.
Why must preprocessing go inside a pipeline?
Imputers, scalers, encoders and selectors all learn from data. Fitted before the split, they carry test information into training. A Pipeline refits every step on each fold’s training rows only. Selecting features outside the loop produced 0.81 AUC on pure noise above, against 0.42 when done correctly.
Key Takeaways
- Audit every column by asking when its value is written in the source system, since anything populated after the outcome inflates your score and predicts nothing.
- Put every imputer, scaler, encoder and selector inside a
Pipeline, because selection outside cross-validation produced 0.81 AUC on data with no signal at all. - Split by time for time-ordered rows and by group when records cluster on a customer or device, rather than defaulting to a random split.
- Deduplicate before splitting, as repeated rows landing on both sides hand the model answers it should have to predict.
- Treat an unexpectedly excellent score as an alarm and investigate it before presenting it, because leakage is the one bug that improves your metrics.