Fewer columns, same answer

Feature selection is the business of removing columns without losing accuracy, and the honest framing is that it usually buys you simplicity rather than performance. On the fraud dataset below, trimming 29 columns to six moved AUC by about 0.002. What it bought was a model with six inputs to maintain instead of twenty-nine.

That is still worth having. Every column in a deployed model is a data pipeline that can break, a schema that can drift, and a dependency on a team that may rename something. Six is cheaper to run than twenty-nine, scores faster, and is far easier to explain.

Three families of method, in increasing order of cost: filters, which score columns independently of any model; wrappers, which repeatedly refit a model; and embedded methods, where the model selects while it trains.

Selection belongs after creating new features rather than instead of it. Generating candidate columns and then pruning them is a reasonable workflow; pruning a thin set of raw columns is just throwing away data. The order matters because selection can only choose among what you built.

Filters: cheap and model-agnostic

Start with the free wins. A column with zero variance carries no information at all.

import numpy as np
import pandas as pd
from sklearn.feature_selection import (VarianceThreshold, SelectKBest,
                                       mutual_info_classif)

# eng, features and y from the feature engineering chapter
noise = pd.DataFrame(rng.normal(0, 1, (m, 20)),
                     columns=[f"noise_{i}" for i in range(20)])
noise["constant"] = 3.0
X = pd.concat([eng[features].reset_index(drop=True), noise], axis=1)
print("columns:", X.shape[1])                                   # 29

keep = VarianceThreshold(threshold=0.0).fit(X)
print("after variance threshold:", int(keep.get_support().sum()))   # 28

scores = pd.Series(mutual_info_classif(X, y, random_state=0), index=X.columns)
print(scores.sort_values(ascending=False).head(5).round(4).to_dict())
# {'hour': 0.0699, 'is_night': 0.0697, 'credit_limit': 0.0079,
#  'noise_3': 0.0058, 'noise_4': 0.0054}

Mutual information measures how much knowing a column reduces uncertainty about the target, and unlike correlation it detects non-linear relationships. That matters whenever a feature has a threshold effect or a U-shaped relationship, both of which score near zero on a correlation coefficient while carrying real signal. It correctly ranks hour and is_night far above everything else.

Note what sits in third and fourth place. Two pure noise columns score 0.0058 and 0.0054, above several real features. With 4000 rows and twenty random columns, some will look mildly informative by chance. Never read a filter ranking as truth; read it as a shortlist.

Method What it measures Cost Catches non-linear
VarianceThreshold Spread within a column Trivial Not applicable
Correlation with target Linear association Trivial No
mutual_info_classif Information shared with target Low Yes
RFE Model performance with each column removed High Depends on model
SelectFromModel with L1 Coefficients the penalty did not zero Low No

What feature selection actually buys

from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.feature_selection import RFE, SelectFromModel
from sklearn.pipeline import make_pipeline
from sklearn.model_selection import cross_val_score

def auc(pipe):
    return round(cross_val_score(pipe, X, y, cv=5, scoring="roc_auc").mean(), 4)

base = LogisticRegression(max_iter=3000)

print("all 29 columns:", auc(make_pipeline(StandardScaler(), base)))
# all 29 columns: 0.9060

print("SelectKBest k=6:", auc(make_pipeline(
    StandardScaler(), SelectKBest(mutual_info_classif, k=6), base)))
# SelectKBest k=6: 0.9074

print("RFE to 6:", auc(make_pipeline(
    StandardScaler(), RFE(LogisticRegression(max_iter=3000), n_features_to_select=6), base)))
# RFE to 6: 0.9079

print("L1 selection:", auc(make_pipeline(
    StandardScaler(),
    SelectFromModel(LogisticRegression(solver="liblinear", l1_ratio=1, C=0.1)),
    base)))
# L1 selection: 0.9052

Four approaches, a spread of 0.0027 AUC between best and worst. Recursive feature elimination refits the model once per removed column and came first by a margin indistinguishable from noise.

So choose on cost rather than score. SelectKBest with mutual information is fast and gets you most of the way. RFE is defensible when you have few columns and time to spare. L1 selection is nearly free because the model does it while fitting.

L1 deserves a caveat, since it is often recommended uncritically. Inspecting what it kept on this data:

chosen = SelectFromModel(
    LogisticRegression(solver="liblinear", l1_ratio=1, C=0.1)
).fit(StandardScaler().fit_transform(X), y)
print(list(X.columns[chosen.get_support()]))
# ['amount_to_limit', 'is_night', 'noise_5', 'noise_8', 'noise_9',
#  'noise_12', 'noise_13', 'noise_17', 'noise_18']

It found both real signals and kept seven noise columns alongside them. The L1 penalty drives coefficients to zero, as the L1 and L2 norms describe, but how aggressively depends entirely on C, and the default is rarely the right value. Tune it or do not trust the result.

When selection matters, and when it does not

Three situations where it genuinely pays.

Linear models with more columns than rows. The fit becomes unstable and selection is not optional.

Deployment cost. Latency, pipeline maintenance and schema dependencies all scale with column count.

Explainability. A six-feature model can be read and defended, using the coefficients described in reading a model’s output. A twenty-nine-feature one cannot be, in practice.

And one situation where it matters less than people assume. Tree ensembles perform their own implicit selection by rarely splitting on uninformative columns:

from sklearn.ensemble import RandomForestClassifier

rf = RandomForestClassifier(n_estimators=200, random_state=0)
print("forest, all 29:", round(cross_val_score(rf, X, y, cv=5, scoring="roc_auc").mean(), 4))
print("forest, 8 good:", round(cross_val_score(rf, eng[features], y, cv=5,
                                               scoring="roc_auc").mean(), 4))
# forest, all 29: 0.9018
# forest, 8 good: 0.8896

The forest given twenty-one junk columns scored slightly higher than the forest given only the good ones. The difference is noise, which is the point: twenty-one useless columns cost it nothing measurable. If you are running gradient boosting on a few dozen features, selection is an operational decision rather than an accuracy one.

One rule overrides everything in this chapter. Selection must happen inside the cross-validation loop, never before it. Choosing columns by looking at the whole dataset uses the target to pick features and then scores on rows that influenced the choice, which inflates results dramatically. Put every selector inside a Pipeline, and see the next chapter for exactly how large that inflation is. Mechanics are documented under RFE.

Frequently Asked Questions

Does feature selection improve model accuracy?

Rarely by much. On the dataset here, cutting 29 columns to six changed AUC by about 0.002. The real gains are operational: fewer pipelines to maintain, faster scoring, and a model you can explain. Accuracy gains appear mainly when columns outnumber rows.

What is the difference between filter, wrapper and embedded methods?

Filters score each column against the target independently of any model, which is fast. Wrappers such as RFE refit a model repeatedly while removing columns, which is slow but model-aware. Embedded methods, such as an L1 penalty, select during training at almost no extra cost.

Do random forests need feature selection?

Not for accuracy. Trees rarely split on uninformative columns, so junk features cost little: adding 21 noise columns here moved a forest by 0.012 AUC, which is within run-to-run variation. Select for deployment simplicity and explainability instead.

Key Takeaways

  • Expect feature selection to buy simplicity rather than accuracy, because six columns scored within 0.003 AUC of twenty-nine on the same data.
  • Start with VarianceThreshold and mutual information, since they cost almost nothing and produce a shortlist worth inspecting.
  • Treat any filter ranking as a shortlist, as two pure noise columns outranked several genuine features here.
  • Tune C before trusting L1 selection, which kept seven noise columns at the value used above.
  • Place every selector inside the pipeline so it refits per fold, never selecting columns on the full dataset before cross-validation.