Four regression metrics, four different questions

A rent model is off by an average of 3,281 rupees. It also has a root mean squared error of 4,110, an R squared of 0.924 and a mean absolute percentage error of 7.9%. All four describe the same predictions. Regression metrics are not interchangeable summaries of quality; each answers a specific question, and picking the wrong one means optimising for something your business does not care about.

The short version. Use MAE when every rupee of error costs the same. Use RMSE when one large miss is worse than several small ones. Use R squared to compare against predicting the average. Use MAPE only when values never approach zero.

import numpy as np
from sklearn.metrics import (mean_absolute_error, root_mean_squared_error,
                             r2_score, mean_absolute_percentage_error)

rng = np.random.default_rng(5)
actual = rng.gamma(3, 9000, 300) + 20000       # monthly rent in rupees
pred = actual + rng.normal(0, 4000, 300)

print("MAE ", round(mean_absolute_error(actual, pred), 1))            # 3281.5
print("RMSE", round(root_mean_squared_error(actual, pred), 1))        # 4110.5
print("R2  ", round(r2_score(actual, pred), 4))                       # 0.9239
print("MAPE", round(mean_absolute_percentage_error(actual, pred), 4)) # 0.0789

MAE and RMSE diverge when one error is large

Mean absolute error averages the size of the misses. Root mean squared error squares them first, averages, then takes the square root, which is why it is always at least as large as MAE and grows faster when errors are uneven.

pred_with_blunder = pred.copy()
pred_with_blunder[0] += 90000        # one flat badly mispriced

print("MAE ", round(mean_absolute_error(actual, pred_with_blunder), 1))
print("RMSE", round(root_mean_squared_error(actual, pred_with_blunder), 1))
# MAE  3581.5   (was 3281.5)
# RMSE 6662.2   (was 4110.5)

One bad prediction out of 300 moved MAE by 300 rupees and RMSE by 2,552. That is not a defect in either metric. It is the choice you are making.

Ask which is worse for the business: being off by 10,000 on one flat, or by 1,000 on ten flats. If the single large error causes a complaint, a refund or a lost customer, RMSE reflects your costs and you should optimise it. If errors simply add up, MAE does. This is the same cost-asymmetry question that framing a business problem asks before any modelling starts.

One consequence worth knowing: squared error is what ordinary least squares minimises, as linear regression from scratch showed. If you report MAE but fit with squared error, your training objective and your reporting metric disagree. Sometimes that is fine. Sometimes it means switching to a model that optimises absolute error directly.

Mean squared error itself is in squared rupees, which nobody can interpret. Report RMSE instead, which is back in the original units and can be read aloud in a meeting without translation.

R squared is not an accuracy percentage

R squared compares your model against a trivial baseline: always predicting the mean of the target. A value of 0.924 means the model explains about 92% of the variation that a constant prediction would leave unexplained.

It is not a percentage of correct predictions, and it has no floor.

flat = np.full_like(actual, actual.mean())
print(round(r2_score(actual, flat), 4))            # 0.0

worse = np.full_like(actual, actual.mean() * 1.4)
print(round(r2_score(actual, worse), 4))           # -1.6014

Predicting the mean every time scores exactly zero, by definition. Predicting 40% above the mean scores -1.6. Negative R squared means your model is worse than a constant, which usually indicates a bug, a broken pipeline, or evaluation on data from a different distribution than training.

Two further cautions. R squared on training data rises whenever you add a column, as multiple linear regression demonstrated, so only the held-out value means anything. And a high R squared on a target with huge natural variation can still leave errors too large to act on, which is why you report it alongside RMSE rather than instead of it.

MAPE and the division problem

Mean absolute percentage error expresses each error as a fraction of the true value, which makes it readable by non-technical stakeholders and dangerous on the wrong data.

small = np.array([0.5, 2.0, 50.0, 1000.0])
preds = np.array([1.0, 2.5, 52.0, 1010.0])

print(np.abs(preds - small))                                     # [ 0.5  0.5  2.  10. ]
print((np.abs(preds - small) / small).round(3))                  # [1.    0.25 0.04 0.01]
print(round(mean_absolute_percentage_error(small, preds), 4))    # 0.325

The first two rows have identical absolute errors of 0.5. One counts as 100% error, the other as 25%, purely because of what they are divided by. The overall MAPE of 32.5% is dominated by the smallest value in the dataset.

Avoid MAPE on anything that can approach zero: demand for slow-moving items, counts, differences. It is also asymmetric, penalising over-prediction more than under-prediction, which quietly biases any model tuned on it towards forecasting low.

Metric Units Punishes large errors Breaks when
MAE Same as target Linearly Rarely
RMSE Same as target Quadratically Outliers you do not care about
R squared Unitless Quadratically Target variance is small or shifting
MAPE Percentage Linearly Any true value near zero

What to report

Report RMSE or MAE as your headline, chosen by what errors actually cost, and R squared beside it for context against the naive baseline. Add a bootstrap interval from the statistics you actually need so nobody argues about a difference of 0.003.

And check your errors against outliers before choosing. If a handful of extreme rows are driving RMSE, the decision about whether to cap them from detecting and treating outliers changes which metric is honest. Definitions are in the r2_score documentation.

Frequently Asked Questions

Should you use MAE or RMSE for regression?

Use RMSE when one large error costs more than several small ones, since squaring makes big misses dominate. Use MAE when errors simply accumulate. In the example above, a single bad prediction out of 300 moved MAE by 300 and RMSE by over 2,500, which is the whole distinction.

What does a negative R squared mean?

Your model performs worse than predicting the mean of the target for every row. Predicting the mean scores exactly zero by definition, so anything below that is a signal of a bug, a broken preprocessing step, or evaluation on data from a different distribution than the training set.

Why is MAPE a bad metric for some problems?

It divides each error by the true value, so rows near zero produce enormous percentages that dominate the average. It is also asymmetric, charging more for over-prediction than under-prediction, which biases models tuned on it towards forecasting low. Avoid it for counts and intermittent demand.

Key Takeaways

  • Choose between MAE and RMSE from what a large error costs the business, rather than reporting whichever looks better.
  • Report RMSE rather than MSE, since squared units cannot be interpreted and RMSE is back on the scale of your target.
  • Read R squared as a comparison against predicting the mean, and treat any negative value as a bug to investigate rather than a weak model.
  • Avoid MAPE whenever the target can approach zero, because one small true value can dominate the entire average.
  • Quote a bootstrap interval alongside any metric so small differences between models are not mistaken for improvements.