Charging the model for confidence

Ordinary least squares has one instruction: minimise squared error on the training data. Give it twenty columns and 280 rows and it will use all twenty, including fifteen that are pure noise, because using them lowers training error slightly. Ridge and lasso regression add a second instruction: keep the coefficients small.

The loss becomes two terms. The first is the usual squared error. The second is a penalty proportional to coefficient size, controlled by a number called alpha. Large coefficients now have to earn their place by reducing error enough to pay the penalty they incur.

The statistical account of regularised regression covers why this reduces variance and what it does to inference. Here the concern is what alpha actually does to your coefficients, which of the two penalties to pick, and the preprocessing step that is genuinely mandatory rather than advisable.

What alpha controls

import numpy as np
import pandas as pd
from sklearn.linear_model import LinearRegression, Ridge, Lasso
from sklearn.preprocessing import StandardScaler
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(19)
N = 400
sqft = rng.normal(1100, 250, N)
X = pd.DataFrame({
    "area_sqft": sqft,
    "area_sqm": sqft * 0.0929 + rng.normal(0, 2, N),     # near-duplicate column
    "rooms": np.clip((sqft / 420 + rng.normal(0, .5, N)).round(), 1, 5),
    "age": rng.integers(0, 35, N).astype(float),
    "metro_km": rng.gamma(2, 1.4, N).clip(.2, 12),
})
for i in range(15):
    X[f"junk_{i}"] = rng.normal(0, 1, N)                 # 15 noise columns

y = (6000 + 21.5 * sqft + 2400 * X["rooms"] - 180 * X["age"]
     - 1500 * X["metro_km"] + rng.normal(0, 4000, N))

X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=.3, random_state=0)
scaler = StandardScaler().fit(X_tr)
X_tr_s, X_te_s = scaler.transform(X_tr), scaler.transform(X_te)

ols = LinearRegression().fit(X_tr_s, y_tr)
print(f"OLS      test R2 {ols.score(X_te_s, y_te):.4f}  "
      f"sum|coef| {np.abs(ols.coef_).sum():9.1f}")
# OLS      test R2 0.7800  sum|coef|   13912.1

for a in [1, 10, 100, 1000]:
    r = Ridge(alpha=a).fit(X_tr_s, y_tr)
    print(f"ridge a={a:<5} test R2 {r.score(X_te_s, y_te):.4f}  "
          f"sum|coef| {np.abs(r.coef_).sum():9.1f}")
# ridge a=1     test R2 0.7795  sum|coef|   13876.0
# ridge a=10    test R2 0.7795  sum|coef|   13628.5
# ridge a=100   test R2 0.7610  sum|coef|   11661.5
# ridge a=1000  test R2 0.4846  sum|coef|    5335.8

alpha is a dial between two failure modes. At zero you have ordinary least squares, free to overfit. At 1000 the penalty dominates, every coefficient is crushed towards zero, and test performance collapses to 0.48. The useful range here is somewhere below 100, and finding it is a tuning problem rather than a judgement call.

Notice that ridge at alpha=1 barely differs from OLS on test performance. On 280 training rows with twenty columns, there is not enough overfitting to fix. The benefit grows as columns approach or exceed rows.

Ridge and lasso regression differ in one way that matters

Ridge penalises the sum of squared coefficients. Lasso penalises the sum of absolute values. That sounds like a technicality and produces completely different behaviour, for the reason set out in the L1 and L2 norms: the absolute-value penalty treats one large coefficient as cheaper than several moderate ones, so it drives the moderate ones to exactly zero.

for a in [1, 50, 200, 800]:
    l = Lasso(alpha=a, max_iter=50000).fit(X_tr_s, y_tr)
    print(f"lasso a={a:<5} test R2 {l.score(X_te_s, y_te):.4f}  "
          f"nonzero {int((l.coef_ != 0).sum())}/20")
# lasso a=1     test R2 0.7801  nonzero 20/20
# lasso a=50    test R2 0.7824  nonzero 16/20
# lasso a=200   test R2 0.7862  nonzero 10/20
# lasso a=800   test R2 0.7601  nonzero 5/20

At alpha=200 the model uses ten of twenty columns and scores slightly better than OLS on all twenty. Ridge never does this; it shrinks every coefficient towards zero but leaves them all nonzero.

Inspect what lasso kept and the picture is less tidy than the usual description:

l200 = Lasso(alpha=200, max_iter=50000).fit(X_tr_s, y_tr)
kept = pd.Series(l200.coef_, index=X.columns)
print(kept[kept != 0].round(1).to_dict())
# {'area_sqft': 3331.6, 'area_sqm': 2472.3, 'rooms': 949.6, 'age': -1486.6,
#  'metro_km': -2672.5, 'junk_6': 121.0, 'junk_7': -64.5, 'junk_8': 231.8,
#  'junk_9': -59.2, 'junk_13': -299.0}

All five genuine features survived, and so did five of the fifteen noise columns. Lasso is a useful filter, not a reliable one, which is the same caution that applied to L1 selection in feature selection techniques.

The two penalties also behave differently on the duplicated area columns. Ridge splits the effect almost evenly, at 2380.1 and 2371.0. Lasso puts 3331.6 on one and 2472.3 on the other, and at higher alpha it would drop one entirely. That is the practical difference when you have correlated columns like those discussed in multiple linear regression: ridge keeps correlated groups together, lasso picks among them somewhat arbitrarily.

Ridge Lasso
Penalty Sum of squared coefficients Sum of absolute coefficients
Coefficients Shrunk, never zero Many set to exactly zero
Correlated columns Effect shared between them One kept, others dropped
Use when All features plausibly matter You want a smaller model
scikit-learn Ridge, RidgeCV Lasso, LassoCV

Scaling first is mandatory, not advisory

The penalty applies to raw coefficient magnitudes, and a coefficient’s magnitude depends on the units of its column. Fit ridge on unscaled data and the penalty lands unevenly.

ols_raw = LinearRegression().fit(X_tr, y_tr)
ridge_raw = Ridge(alpha=100).fit(X_tr, y_tr)

for col in ["area_sqft", "rooms", "age", "metro_km"]:
    j = list(X.columns).index(col)
    print(f"{col:10s} OLS {ols_raw.coef_[j]:9.2f} -> ridge {ridge_raw.coef_[j]:9.2f}"
          f"   kept {ridge_raw.coef_[j] / ols_raw.coef_[j]:6.1%}")
# area_sqft  OLS     16.50 -> ridge     18.08   kept 109.6%
# rooms      OLS   1285.19 -> ridge    578.23   kept  45.0%
# age        OLS   -161.77 -> ridge   -164.77   kept 101.9%
# metro_km   OLS  -1372.88 -> ridge  -1277.42   kept  93.0%

rooms lost 55% of its coefficient. area_sqft was left completely alone. Neither outcome has anything to do with which feature matters; it is entirely about units. Area runs into the thousands so its coefficient is small and cheap to keep, while rooms runs from 1 to 5 so its coefficient is large and expensive.

Put a StandardScaler before any penalised model, as covered in feature scaling and normalisation, and inside a pipeline so it refits per fold.

Picking alpha, and whether elastic net adds anything

Cross-validated variants do the search for you.

from sklearn.linear_model import RidgeCV, LassoCV, ElasticNetCV

rcv = RidgeCV(alphas=np.logspace(-2, 4, 40)).fit(X_tr_s, y_tr)
print(round(float(rcv.alpha_), 2), round(rcv.score(X_te_s, y_te), 4))
# 17.01 0.7794

lcv = LassoCV(alphas=np.logspace(-1, 3.5, 40), max_iter=50000,
              random_state=0).fit(X_tr_s, y_tr)
print(round(float(lcv.alpha_), 2), round(lcv.score(X_te_s, y_te), 4),
      int((lcv.coef_ != 0).sum()))
# 289.43 0.7855 8

encv = ElasticNetCV(l1_ratio=[.5, .7, .9, .95, .99, 1.0],
                    alphas=np.logspace(-1, 3.5, 40),
                    max_iter=100000, cv=5, random_state=0).fit(X_tr_s, y_tr)
print(encv.l1_ratio_, round(float(encv.alpha_), 2), round(encv.score(X_te_s, y_te), 4))
# 1.0 289.43 0.7855

Elastic net blends both penalties, and here cross-validation selected l1_ratio=1.0, which is pure lasso. The blend added nothing on this data. That happens often enough that I would reach for LassoCV or RidgeCV first and try elastic net only when you have groups of correlated columns and want to keep whole groups rather than one member. Parameter details are in the Ridge and Lasso references.

Frequently Asked Questions

What is the difference between ridge and lasso regression?

Ridge penalises squared coefficients and shrinks all of them towards zero without eliminating any. Lasso penalises absolute values and sets many to exactly zero, producing a smaller model. With correlated columns, ridge shares the effect between them while lasso keeps one and drops the rest.

What does the alpha parameter do in ridge regression?

It sets how heavily coefficient size is penalised. At zero you get ordinary least squares. As alpha rises, coefficients shrink and the model becomes more stable but less flexible. At alpha 1000 in the example above, test R squared fell from 0.78 to 0.48. Select it by cross-validation.

Do you need to scale features before ridge or lasso?

Yes, always. The penalty acts on raw coefficient magnitudes, so a column measured in thousands gets a small coefficient and escapes the penalty while a column measured in single digits gets crushed. Unscaled ridge here removed 55% of the rooms coefficient and left area untouched.

Key Takeaways

  • Put a StandardScaler before every penalised model, because without it the penalty falls on columns according to their units rather than their importance.
  • Select alpha with RidgeCV or LassoCV rather than guessing, since performance collapsed from 0.78 to 0.48 between reasonable-looking values.
  • Choose lasso when you want a smaller model and ridge when all features plausibly matter, particularly where correlated columns should stay together.
  • Treat lasso’s selection as a filter rather than a verdict, as it retained five of fifteen pure noise columns here.
  • Try elastic net only for grouped correlated features, because cross-validation chose pure lasso over any blend on this dataset.