An outlier is a judgement, not a measurement

No test tells you a value is an outlier. Every method for detecting outliers applies a threshold someone chose, to a definition of “far from typical” someone else chose, and different reasonable choices flag wildly different numbers of rows in the same column.

Here is what that looks like on the insurance mileage column flagged during exploratory data analysis. A z-score rule flags 18 rows. The interquartile range rule flags 69. A robust z-score flags

  1. Same 1265 values, three defensible methods, a near four-fold spread.

So the useful question is not “is this an outlier” but “will this value distort the model I am about to fit”. Sometimes the answer is yes and you act. Often the answer is no, and the extreme value is the most interesting record in your dataset.

Three methods for detecting outliers, and why they disagree

import numpy as np
import pandas as pd

mileage = data["annual_mileage"].dropna()      # 1265 values, profiled earlier

z = (mileage - mileage.mean()) / mileage.std()

q1, q3 = mileage.quantile([0.25, 0.75])
iqr = q3 - q1
low, high = q1 - 1.5 * iqr, q3 + 1.5 * iqr

median = mileage.median()
mad = (mileage - median).abs().median()         # median absolute deviation
robust_z = 0.6745 * (mileage - median) / mad

print("z > 3:          ", int((z.abs() > 3).sum()))                      # 18
print("IQR fence:      ", round(low, 1), round(high, 1),
      "flagged:", int(((mileage < low) | (mileage > high)).sum()))
# IQR fence: -10476.5 33052.2 flagged: 69
print("robust z > 3.5: ", int((robust_z.abs() > 3.5).sum()))             # 54
print("skew:           ", round(float(mileage.skew()), 2))               # 9.48
print("max:", round(float(mileage.max()), 1),
      "p99:", round(float(mileage.quantile(0.99)), 1))
# max: 396751.2 p99: 87855.8

The skew of 9.48 explains the disagreement. Z-scores use the mean and standard deviation, and both are dragged upward by the extreme values you are trying to find, so the threshold moves out to meet them and the method flags fewer rows than it should. This is the masking problem, and it gets worse the more outliers you have.

The IQR rule uses quartiles, which extreme values cannot move, so it keeps working on skewed data. The robust z-score uses the median and median absolute deviation for the same reason.

Method Based on Good for Fails when
Z-score, threshold 3 Mean, standard deviation Roughly symmetric data Skewed data, or several outliers masking each other
IQR, 1.5 fence Quartiles Skewed data, quick default Flags a lot on heavy-tailed distributions
Robust z, threshold 3.5 Median, MAD Skewed data, several outliers Breaks if more than half the values are identical
Isolation Forest Multivariate structure Combinations that are odd together Single-column questions, where it is overkill

On any skewed column, which covers most monetary and count data, use the IQR fence or the robust z-score. Z-scores on skewed data are the most common mistake in this area, and they fail silently by under-reporting.

The note on the fourth row matters when outliers are only strange in combination. A driver aged 23 is ordinary and 40,000 annual miles is ordinary, but together they may be rare. Single-column rules cannot see that, and Section 9 covers the multivariate approach properly.

Deciding what to actually do

Four options, and the choice depends on where the value came from.

Leave it. The default, and correct more often than people expect. A genuine customer with 390,000 miles is a real customer. Removing them because the number is inconvenient means building a model that fails on exactly the cases where mistakes cost most.

Fix it. When you can establish the value is an error, correct or blank it. A driver aged 999 is a form default, not a person. Set it to NaN and let your missing value strategy handle it.

Cap it. Winsorising clips values at a percentile, keeping the row and its other columns while limiting the extreme’s influence. This is my usual choice when the values are real but the tail is distorting a linear model.

Drop the row. Only when the record is clearly corrupt across several columns. Dropping rows because one value is large discards real information and biases the model towards the middle of your distribution.

Transforming the column with a logarithm is a fifth route that often removes the problem entirely, since it compresses a long right tail into something close to symmetric.

Who actually cares about outliers

The answer depends almost entirely on your model family.

from sklearn.linear_model import LinearRegression
from sklearn.ensemble import RandomForestRegressor
from sklearn.model_selection import train_test_split
from sklearn.metrics import mean_absolute_error

rng = np.random.default_rng(3)
x = rng.normal(50, 10, 600)
y_true = 3 * x + rng.normal(0, 15, 600)
x[:6] += 400                                    # six corrupted sensor readings
X_raw = x.reshape(-1, 1)

for label, Xv in [("raw", X_raw),
                  ("capped at p99", np.clip(X_raw, None, np.percentile(X_raw, 99)))]:
    Xtr, Xte, ytr, yte = train_test_split(Xv, y_true, test_size=0.3, random_state=0)
    lin = LinearRegression().fit(Xtr, ytr)
    rf = RandomForestRegressor(n_estimators=200, random_state=0).fit(Xtr, ytr)
    print(label,
          "| linear MAE", round(mean_absolute_error(yte, lin.predict(Xte)), 2),
          "| forest MAE", round(mean_absolute_error(yte, rf.predict(Xte)), 2))

# raw           | linear MAE 26.77 | forest MAE 15.45
# capped at p99 | linear MAE 13.92 | forest MAE 15.52

Capping halves the linear model’s error, from 26.77 to 13.92. The random forest moves from 15.45 to 15.52, which is noise.

That difference is the whole practical story. Linear regression minimises squared error, so one point a long way out pulls the entire fitted line towards it. Trees split on ordering rather than distance, so an extreme value simply sits in the rightmost leaf and changes nothing else. Distance-based methods such as k-nearest neighbours and support vector machines sit with the linear models in sensitivity.

Which means outlier treatment is a decision about your model, not about your data. If you are fitting gradient boosting, you can often skip this chapter’s treatment step entirely and keep the information. If you are fitting linear or distance-based models, cap the tails. Percentile mechanics are documented under numpy.percentile, and the sampling considerations behind quantile estimates appear in the statistics you actually need.

Whatever you do, do it inside the pipeline and record the thresholds, because a cap learned from the full dataset is another route for test information to reach training.

Frequently Asked Questions

Should you remove outliers before training a model?

Usually not. Remove a value only when you can establish it is an error. Genuine extreme values are often the cases that matter most, and deleting rows biases your model towards the middle of the distribution. Capping at a percentile keeps the record while limiting its influence.

Is the z-score or IQR method better for detecting outliers?

IQR on skewed data, which covers most monetary and count columns. Z-scores rely on the mean and standard deviation, which the outliers themselves inflate, so the method under-reports. On the mileage column above, z-scores flagged 18 rows where the IQR rule flagged 69.

Do outliers affect random forests?

Barely. Trees split on the ordering of values rather than their distance, so an extreme reading falls into the end leaf and changes nothing else. In the comparison above, capping cut linear regression error by half and moved the forest by 0.07, which is noise.

Key Takeaways

  • Use the IQR fence or a robust z-score on any skewed column, because standard z-scores are inflated by the very values they are meant to detect.
  • Decide treatment from your model family, since capping halved linear regression error and left a random forest unchanged on the same data.
  • Cap at a percentile rather than dropping rows, so you keep the record’s other columns and avoid biasing the model towards the middle.
  • Convert values you can prove are errors into NaN and let your imputation strategy handle them, rather than inventing a replacement.
  • Fit any cap or threshold on training data only, and record the exact values so the same treatment applies at prediction time.