Columns with different units are not comparable
A credit card balance runs to seven figures. Utilisation is a fraction between 0 and 1. Measure the distance between two customers across both columns and the balance decides the answer entirely, because a difference of 40,000 rupees swamps a difference of 0.4 in utilisation. Feature scaling removes that accident of units so every column gets a fair hearing.
You have already seen the consequence: in your first model in scikit-learn, standardising moved a k-nearest neighbours classifier from 0.667 to 0.933 on identical data. This chapter is about which scaler to pick, what each one does to a skewed column, and which models can safely skip the step.
The statistical treatment of standardisation and normalisation covers the distributional reasoning. Here the question is which transformer to put in your pipeline.
Three scalers, one skewed column
import numpy as np
import pandas as pd
from sklearn.preprocessing import StandardScaler, MinMaxScaler, RobustScaler
rng = np.random.default_rng(21)
n = 2000
cards = pd.DataFrame({
"age": rng.integers(21, 78, n).astype(float),
"balance": rng.gamma(1.5, 42000, n),
"utilisation": rng.beta(2, 5, n),
})
cards.loc[rng.choice(n, 12, replace=False), "balance"] *= 14 # 12 extreme balances
print(round(float(cards["balance"].skew()), 2),
round(float(cards["balance"].median())),
round(float(cards["balance"].max())))
# 11.6 48886 2142446
for name, scaler in [("standard", StandardScaler()),
("minmax", MinMaxScaler()),
("robust", RobustScaler())]:
b = scaler.fit_transform(cards)[:, 1]
print(f"{name:9s} median {np.median(b):7.3f} p99 {np.percentile(b, 99):7.3f}"
f" max {b.max():8.3f}")
# standard median -0.167 p99 1.788 max 18.122
# minmax median 0.023 p99 0.127 max 1.000
# robust median 0.000 p99 3.805 max 35.600
Look at the minmax row. Twelve extreme balances out of 2000 have pushed 99% of the data below 0.127, so almost every customer is now crammed into the bottom eighth of the range. The column is technically scaled and practically useless.
StandardScaler subtracts the mean and divides by the standard deviation, which produces a column centred near zero with unit spread. MinMaxScaler squeezes everything into 0 to 1 using the observed minimum and maximum, which is why a single extreme value controls the result. RobustScaler uses the median and the interquartile range instead, so the outliers you identified in detecting and treating outliers cannot move the transformation.
| Scaler | Centres on | Divides by | Output range | Outlier resistant |
|---|---|---|---|---|
StandardScaler |
Mean | Standard deviation | Unbounded, mostly -3 to 3 | No |
MinMaxScaler |
Minimum | Range | 0 to 1 on training data | No, badly |
RobustScaler |
Median | Interquartile range | Unbounded | Yes |
MaxAbsScaler |
Nothing | Largest absolute value | -1 to 1 | No |
Use StandardScaler as your default. Switch to RobustScaler when a column is skewed or carries genuine extreme values. Reach for MinMaxScaler only when something downstream requires a bounded input, such as an image pipeline or a neural network layer expecting 0 to 1.
MinMaxScaler breaks outside its training range
from sklearn.model_selection import train_test_split
train, test = train_test_split(cards, test_size=0.3, random_state=0)
scaled_test = MinMaxScaler().fit(train).transform(test)[:, 1]
print(round(float(scaled_test.min()), 3), round(float(scaled_test.max()), 3))
# 0.001 1.018
print("values outside [0, 1]:", int(((scaled_test < 0) | (scaled_test > 1)).sum()))
# values outside [0, 1]: 1
The guaranteed 0 to 1 range is a guarantee about training data only. One test customer had a balance above anything seen in training, so it scaled to 1.018. Harmless here, fatal if a downstream component assumes the bound holds.
Which models actually need feature scaling
from sklearn.neighbors import KNeighborsClassifier
from sklearn.svm import SVC
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
from sklearn.pipeline import make_pipeline
from sklearn.model_selection import cross_val_score
risk = (0.000012 * cards["balance"] + 3.2 * cards["utilisation"]
- 0.02 * cards["age"])
y = (risk + rng.normal(0, 0.4, n) > 0.9).astype(int)
for name, model in [("KNN", KNeighborsClassifier()),
("SVC", SVC()),
("LogReg", LogisticRegression(max_iter=3000)),
("RandomForest", RandomForestClassifier(n_estimators=150,
random_state=0))]:
raw = cross_val_score(model, cards, y, cv=5, scoring="roc_auc").mean()
scaled = cross_val_score(make_pipeline(StandardScaler(), model), cards, y,
cv=5, scoring="roc_auc").mean()
print(f"{name:13s} raw {raw:.4f} scaled {scaled:.4f} delta {scaled - raw:+.4f}")
# KNN raw 0.7313 scaled 0.9205 delta +0.1893
# SVC raw 0.7550 scaled 0.9487 delta +0.1937
# LogReg raw 0.9499 scaled 0.9505 delta +0.0006
# RandomForest raw 0.9331 scaled 0.9330 delta -0.0000
Two models gain roughly 0.19 of AUC. Two gain nothing.
The split follows from how each algorithm uses the numbers. K-nearest neighbours and support vector machines measure distance between rows, so unit differences distort similarity directly. Trees split one column at a time on an ordering, so multiplying a column by a thousand changes nothing about where the splits fall.
Logistic regression sits in between, and the near-zero delta here is slightly misleading. The fit converges to the same answer either way, but it converges faster when scaled, and the coefficients only become comparable after scaling, as reading a model’s output showed. Any model with an L1 or L2 penalty genuinely does need scaling, because the penalty is applied to raw coefficient sizes.
| Model family | Scaling needed | Why |
|---|---|---|
| k-nearest neighbours, SVM, k-means | Yes, always | They measure distance across columns |
| Linear and logistic regression, unpenalised | Optional | Same fit, faster convergence, readable coefficients |
| Ridge, lasso, elastic net | Yes | The penalty acts on raw coefficient magnitudes |
| Neural networks | Yes | Gradient descent converges badly on mixed scales |
| Decision trees, random forests, boosting | No | Splits depend on order, not distance |
The practical rule: scale by default. It costs one line, it is required by several families, and it is harmless for the rest.
Fit the scaler on training folds only. A scaler learns a mean or a median from data, so fitting it on everything puts test information into training, which is why it belongs inside a Pipeline rather than applied to your dataframe in a notebook. Parameter details are in the StandardScaler and RobustScaler references.
Frequently Asked Questions
What is the difference between normalisation and standardisation?
Standardisation subtracts the mean and divides by the standard deviation, giving a column centred on zero with unit spread and no fixed bounds. Normalisation, usually meaning min-max scaling, compresses values into 0 to 1 using the observed minimum and maximum, which a single extreme value can dominate.
Do decision trees need feature scaling?
No. Trees and tree ensembles split on the ordering of values within one column at a time, so multiplying a column by any constant produces exactly the same splits. In the comparison above, scaling moved a random forest by less than 0.0001 AUC while moving k-nearest neighbours by 0.19.
Should you scale before or after the train test split?
After, and inside a pipeline. The scaler learns a mean or median from the data it is fitted on, so fitting it on the full dataset lets test rows influence the transformation applied during training. Putting StandardScaler in a Pipeline refits it correctly on every fold.
Key Takeaways
- Default to
StandardScalerand switch toRobustScalerwhenever a column is skewed, because min-max scaling let twelve extreme values push 99% of the data below 0.127. - Always scale for k-nearest neighbours, support vector machines and any penalised linear model, since those two families gained about 0.19 AUC here.
- Skip scaling for trees and tree ensembles without guilt, as their splits depend on ordering rather than distance.
- Avoid
MinMaxScalerwhen anything downstream assumes a hard 0 to 1 bound, because unseen values exceed the training range in production. - Put the scaler inside a
Pipelineso it refits on each training fold instead of learning statistics from rows you intend to test on.