Why it is blank decides what to do about it
Most scikit-learn estimators refuse to fit on NaN, so handling missing values is not optional. The question is never whether to deal with them but what the blank means, because a sensor that failed and a customer who declined to answer require different treatment and only one of them is really missing.
The practical answer, up front: median-impute inside a pipeline, add an indicator column when the blank looks informative, and only drop rows when very few are affected. Everything else in this chapter is about knowing when that default is wrong.
The statistical treatment of imputation and deletion covers the inference side. Here the concern is what each choice does to a model’s score and where in your code it has to happen.
Three mechanisms, three consequences
| Mechanism | Meaning | Example | Consequence |
|---|---|---|---|
| MCAR | Missing completely at random | A form page timed out at random | Imputation is safe, deletion is unbiased |
| MAR | Missingness depends on other observed columns | Older drivers skip the mileage question | Imputation works if those columns are in the model |
| MNAR | Missingness depends on the unseen value itself | High earners decline to state income | Imputation biases the result, and an indicator is essential |
The distinction is not academic. The mileage column from exploratory data analysis is missing mostly for drivers over 65, which is MAR. Replacing those blanks with the overall median assigns elderly, low-mileage drivers a typical driver’s annual distance, which is wrong in a specific and predictable direction.
You usually cannot prove which mechanism applies. What you can do is check whether missingness correlates with anything else you hold, and assume MNAR when the blank plausibly depends on the value itself.
What handling missing values costs you
Four options, compared on the same data with the same model.
import numpy as np
import pandas as pd
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import HistGradientBoostingClassifier
from sklearn.pipeline import make_pipeline
from sklearn.model_selection import cross_val_score
# df from the previous chapter, duplicates removed
data = df.drop_duplicates()
features = ["driver_age", "annual_mileage", "prior_claims", "survey_score"]
X, y = data[features], data["claimed"]
print(X.isna().mean().round(3).to_dict())
# {'driver_age': 0.0, 'annual_mileage': 0.157,
# 'prior_claims': 0.0, 'survey_score': 0.267}
print("rows kept if you drop any row with a blank:", len(X.dropna()), "of", len(X))
# rows kept if you drop any row with a blank: 928 of 1500
for strategy in ["mean", "median", "most_frequent"]:
pipe = make_pipeline(SimpleImputer(strategy=strategy),
StandardScaler(),
LogisticRegression(max_iter=2000))
print(strategy, round(cross_val_score(pipe, X, y, cv=5, scoring="roc_auc").mean(), 4))
# mean 0.8975
# median 0.8966
# most_frequent 0.8925
Two things to take from that. Dropping incomplete rows would discard 572 of 1500 records, 38% of the data, which is a severe price for a convenience. And the three imputation strategies differ by less than a hundredth of an AUC point.
That second result is typical and worth internalising. People spend hours choosing between mean and median imputation on datasets where the choice changes nothing measurable. Median is the sensible default because it resists skew, and then you should move on to a decision that matters.
Deletion is still right in one case: when under about 5% of rows are affected and the missingness looks random. Below that threshold you lose little and gain simplicity. Column deletion is right when a column is 60% or more missing and carries weak signal, which describes survey_score reasonably well given its 0.084 correlation with the target.
The indicator column, and when it adds nothing
Adding a binary flag for “this value was missing” lets the model use the fact of the blank, which is exactly what you want when missingness is informative.
pipe = make_pipeline(
SimpleImputer(strategy="median", add_indicator=True),
StandardScaler(),
LogisticRegression(max_iter=2000))
print(round(cross_val_score(pipe, X, y, cv=5, scoring="roc_auc").mean(), 4))
# 0.8975
Identical to plain median imputation, to four decimal places. The missingness here is genuinely informative, so why did the flag add nothing?
Because driver_age is already in the model. Mileage goes missing for older drivers, and the model can already see their age, so the indicator is a near-copy of information it holds. An indicator earns its place when the thing that drives the missingness is not otherwise represented in your features.
Keep adding it anyway. It costs one column, it is free when redundant, and it is the only defence you have against MNAR columns where the blank is the signal.
Models that take NaN directly
Some estimators handle missing values natively and need no imputation at all.
hgb = HistGradientBoostingClassifier(random_state=0)
print(round(cross_val_score(hgb, X, y, cv=5, scoring="roc_auc").mean(), 4))
# 0.8597
Lower than the imputed logistic regression, at 0.860 against 0.898. Native handling is a convenience rather than an upgrade, and on a 1500-row dataset a boosted ensemble has less to work with than a linear model does.
HistGradientBoostingClassifier, HistGradientBoostingRegressor, XGBoost and LightGBM all learn a default direction for missing values at each split, which is genuinely useful when missingness is informative and you have enough data. Most other scikit-learn estimators will raise a ValueError on NaN.
Where the imputation has to happen
Imputation learns a statistic from data, so it belongs inside the pipeline.
from sklearn.model_selection import train_test_split
print(SimpleImputer(strategy="mean").fit(X).statistics_.round(2))
# [5.053000e+01 1.406041e+04 5.000000e-01 7.010000e+00]
X_train, X_test = train_test_split(X, test_size=0.3, random_state=0)
print(SimpleImputer(strategy="mean").fit(X_train).statistics_.round(2))
# [5.0540e+01 1.4157e+04 4.9000e-01 7.0000e+00]
The mileage mean is 14060 computed on everything and 14157 computed on the training split alone. The gap is small on 1500 rows and it will not be small on 200. More to the point, the first number contains information from rows you are about to test on, which is the quiet failure mode described in the main challenges of machine learning. Putting SimpleImputer inside a Pipeline makes this impossible to get wrong, since each fold refits the imputer on its own training rows. Full parameter details are in the SimpleImputer documentation.
Frequently Asked Questions
Should you use mean or median imputation?
Median, as a default, because it is unaffected by skew and outliers. The honest answer is that the difference is usually negligible: on the dataset above, mean and median imputation differed by 0.0009 AUC. Spend the time on whether the missingness is informative instead.
When should you drop rows with missing values?
When under roughly 5% of rows are affected and nothing suggests the blanks are systematic. Above that the cost is severe: dropping every incomplete row in the example here would discard 38% of the data, including most of the older drivers who behave differently from everyone else.
Can machine learning models handle missing values automatically?
Histogram-based gradient boosting in scikit-learn, along with XGBoost and LightGBM, learns a default split direction for missing values. Most other estimators raise an error. Native handling is convenient rather than superior, and it scored lower than imputed logistic regression on the example above.
Key Takeaways
- Check whether missingness correlates with your other columns before choosing a treatment, because a blank that depends on the unseen value cannot be imputed honestly.
- Default to median imputation inside a pipeline, and stop deliberating between mean and median when the measured difference is under a hundredth of an AUC point.
- Add
add_indicator=Trueas standard practice, accepting that it contributes nothing when another feature already encodes the same pattern. - Reserve row deletion for cases under about 5% missing, since dropping incomplete records removed 38% of the data in the example here.
- Fit every imputer on training folds only, which a
Pipelineenforces for you and a notebook does not.