Columns with different units are not comparable

A credit card balance runs to seven figures. Utilisation is a fraction between 0 and 1. Measure the distance between two customers across both columns and the balance decides the answer entirely, because a difference of 40,000 rupees swamps a difference of 0.4 in utilisation. Feature scaling removes that accident of units so every column gets a fair hearing.

You have already seen the consequence: in your first model in scikit-learn, standardising moved a k-nearest neighbours classifier from 0.667 to 0.933 on identical data. This chapter is about which scaler to pick, what each one does to a skewed column, and which models can safely skip the step.

The statistical treatment of standardisation and normalisation covers the distributional reasoning. Here the question is which transformer to put in your pipeline.

Three scalers, one skewed column

import numpy as np
import pandas as pd
from sklearn.preprocessing import StandardScaler, MinMaxScaler, RobustScaler

rng = np.random.default_rng(21)
n = 2000
cards = pd.DataFrame({
    "age": rng.integers(21, 78, n).astype(float),
    "balance": rng.gamma(1.5, 42000, n),
    "utilisation": rng.beta(2, 5, n),
})
cards.loc[rng.choice(n, 12, replace=False), "balance"] *= 14   # 12 extreme balances

print(round(float(cards["balance"].skew()), 2),
      round(float(cards["balance"].median())),
      round(float(cards["balance"].max())))
# 11.6 48886 2142446

for name, scaler in [("standard", StandardScaler()),
                     ("minmax", MinMaxScaler()),
                     ("robust", RobustScaler())]:
    b = scaler.fit_transform(cards)[:, 1]
    print(f"{name:9s} median {np.median(b):7.3f}  p99 {np.percentile(b, 99):7.3f}"
          f"  max {b.max():8.3f}")
# standard  median  -0.167  p99   1.788  max   18.122
# minmax    median   0.023  p99   0.127  max    1.000
# robust    median   0.000  p99   3.805  max   35.600

Look at the minmax row. Twelve extreme balances out of 2000 have pushed 99% of the data below 0.127, so almost every customer is now crammed into the bottom eighth of the range. The column is technically scaled and practically useless.

StandardScaler subtracts the mean and divides by the standard deviation, which produces a column centred near zero with unit spread. MinMaxScaler squeezes everything into 0 to 1 using the observed minimum and maximum, which is why a single extreme value controls the result. RobustScaler uses the median and the interquartile range instead, so the outliers you identified in detecting and treating outliers cannot move the transformation.

Scaler Centres on Divides by Output range Outlier resistant
StandardScaler Mean Standard deviation Unbounded, mostly -3 to 3 No
MinMaxScaler Minimum Range 0 to 1 on training data No, badly
RobustScaler Median Interquartile range Unbounded Yes
MaxAbsScaler Nothing Largest absolute value -1 to 1 No

Use StandardScaler as your default. Switch to RobustScaler when a column is skewed or carries genuine extreme values. Reach for MinMaxScaler only when something downstream requires a bounded input, such as an image pipeline or a neural network layer expecting 0 to 1.

MinMaxScaler breaks outside its training range

from sklearn.model_selection import train_test_split

train, test = train_test_split(cards, test_size=0.3, random_state=0)
scaled_test = MinMaxScaler().fit(train).transform(test)[:, 1]

print(round(float(scaled_test.min()), 3), round(float(scaled_test.max()), 3))
# 0.001 1.018
print("values outside [0, 1]:", int(((scaled_test < 0) | (scaled_test > 1)).sum()))
# values outside [0, 1]: 1

The guaranteed 0 to 1 range is a guarantee about training data only. One test customer had a balance above anything seen in training, so it scaled to 1.018. Harmless here, fatal if a downstream component assumes the bound holds.

Which models actually need feature scaling

from sklearn.neighbors import KNeighborsClassifier
from sklearn.svm import SVC
from sklearn.linear_model import LogisticRegression
from sklearn.ensemble import RandomForestClassifier
from sklearn.pipeline import make_pipeline
from sklearn.model_selection import cross_val_score

risk = (0.000012 * cards["balance"] + 3.2 * cards["utilisation"]
        - 0.02 * cards["age"])
y = (risk + rng.normal(0, 0.4, n) > 0.9).astype(int)

for name, model in [("KNN", KNeighborsClassifier()),
                    ("SVC", SVC()),
                    ("LogReg", LogisticRegression(max_iter=3000)),
                    ("RandomForest", RandomForestClassifier(n_estimators=150,
                                                            random_state=0))]:
    raw = cross_val_score(model, cards, y, cv=5, scoring="roc_auc").mean()
    scaled = cross_val_score(make_pipeline(StandardScaler(), model), cards, y,
                             cv=5, scoring="roc_auc").mean()
    print(f"{name:13s} raw {raw:.4f}  scaled {scaled:.4f}  delta {scaled - raw:+.4f}")

# KNN           raw 0.7313  scaled 0.9205  delta +0.1893
# SVC           raw 0.7550  scaled 0.9487  delta +0.1937
# LogReg        raw 0.9499  scaled 0.9505  delta +0.0006
# RandomForest  raw 0.9331  scaled 0.9330  delta -0.0000

Two models gain roughly 0.19 of AUC. Two gain nothing.

The split follows from how each algorithm uses the numbers. K-nearest neighbours and support vector machines measure distance between rows, so unit differences distort similarity directly. Trees split one column at a time on an ordering, so multiplying a column by a thousand changes nothing about where the splits fall.

Logistic regression sits in between, and the near-zero delta here is slightly misleading. The fit converges to the same answer either way, but it converges faster when scaled, and the coefficients only become comparable after scaling, as reading a model’s output showed. Any model with an L1 or L2 penalty genuinely does need scaling, because the penalty is applied to raw coefficient sizes.

Model family Scaling needed Why
k-nearest neighbours, SVM, k-means Yes, always They measure distance across columns
Linear and logistic regression, unpenalised Optional Same fit, faster convergence, readable coefficients
Ridge, lasso, elastic net Yes The penalty acts on raw coefficient magnitudes
Neural networks Yes Gradient descent converges badly on mixed scales
Decision trees, random forests, boosting No Splits depend on order, not distance

The practical rule: scale by default. It costs one line, it is required by several families, and it is harmless for the rest.

Fit the scaler on training folds only. A scaler learns a mean or a median from data, so fitting it on everything puts test information into training, which is why it belongs inside a Pipeline rather than applied to your dataframe in a notebook. Parameter details are in the StandardScaler and RobustScaler references.

Frequently Asked Questions

What is the difference between normalisation and standardisation?

Standardisation subtracts the mean and divides by the standard deviation, giving a column centred on zero with unit spread and no fixed bounds. Normalisation, usually meaning min-max scaling, compresses values into 0 to 1 using the observed minimum and maximum, which a single extreme value can dominate.

Do decision trees need feature scaling?

No. Trees and tree ensembles split on the ordering of values within one column at a time, so multiplying a column by any constant produces exactly the same splits. In the comparison above, scaling moved a random forest by less than 0.0001 AUC while moving k-nearest neighbours by 0.19.

Should you scale before or after the train test split?

After, and inside a pipeline. The scaler learns a mean or median from the data it is fitted on, so fitting it on the full dataset lets test rows influence the transformation applied during training. Putting StandardScaler in a Pipeline refits it correctly on every fold.

Key Takeaways

  • Default to StandardScaler and switch to RobustScaler whenever a column is skewed, because min-max scaling let twelve extreme values push 99% of the data below 0.127.
  • Always scale for k-nearest neighbours, support vector machines and any penalised linear model, since those two families gained about 0.19 AUC here.
  • Skip scaling for trees and tree ensembles without guilt, as their splits depend on ordering rather than distance.
  • Avoid MinMaxScaler when anything downstream assumes a hard 0 to 1 bound, because unseen values exceed the training range in production.
  • Put the scaler inside a Pipeline so it refits on each training fold instead of learning statistics from rows you intend to test on.