The mapping you choose becomes an assumption

Encoding categorical variables is the step where a column of text becomes a column of numbers, and the mapping you pick silently tells the model something about the categories. Number them 0, 1, 2, 3 and you have asserted that West is three times North and that South sits between them. For a linear model, that assertion is arithmetic it will act on.

The default that avoids this is one-hot encoding: one binary column per category, no implied order, no implied distance. Use it unless the categories genuinely have an order or there are too many of them to afford.

The statistical treatment of one-hot, label and target encoding covers why each affects coefficients the way it does. This chapter is about the scikit-learn transformers, the shapes they produce, and what happens when a category appears at prediction time that was not in training.

Encoding categorical variables with one-hot by default

import pandas as pd
from sklearn.preprocessing import OneHotEncoder

cats = data[["region", "policy_tier", "vehicle_make"]]
print(cats.nunique().to_dict())
# {'region': 4, 'policy_tier': 3, 'vehicle_make': 60}

ohe = OneHotEncoder(handle_unknown="ignore", sparse_output=False)
encoded = ohe.fit_transform(cats[["region", "policy_tier"]])

print(encoded.shape)                        # (1500, 7)
print(list(ohe.get_feature_names_out()))
# ['region_East', 'region_North', 'region_South', 'region_West',
#  'policy_tier_bronze', 'policy_tier_gold', 'policy_tier_silver']

Four regions and three tiers became seven columns, each holding 1 or 0. Every category is now equidistant from every other, which is the honest representation of unordered labels, and the dot product a linear model computes assigns each one its own coefficient.

Two arguments are worth setting deliberately. sparse_output=False returns a dense array, which is easier to inspect and fine until you have hundreds of columns. handle_unknown="ignore" is covered below and is the one that prevents a production outage.

There is also drop="first", which removes one column per feature to avoid perfect collinearity between the dummies and the intercept. It matters for the coefficient interpretation that reading a model’s output depends on, and it matters not at all for trees. Leave it off unless you are reading coefficients from an unregularised linear model.

Ordinal encoding, and when it is a mistake

OrdinalEncoder maps categories to integers. It sorts alphabetically unless you tell it otherwise.

from sklearn.preprocessing import OrdinalEncoder

oe = OrdinalEncoder().fit(cats[["region"]])
print(oe.categories_)
# [array(['East', 'North', 'South', 'West'], dtype=object)]
print(oe.transform(pd.DataFrame({"region": ["North", "South", "East", "West"]})).ravel())
# [1. 2. 0. 3.]

East is 0 and West is 3, purely because of the alphabet. A linear model now treats West as three units of whatever North is one unit of, which is meaningless.

Ordinal encoding is correct when the order is real, and policy_tier is the case in point. Bronze, silver and gold have a genuine ranking, so give it explicitly rather than accepting the alphabet, which would produce bronze, gold, silver.

tier = OrdinalEncoder(categories=[["bronze", "silver", "gold"]])
print(tier.fit_transform(pd.DataFrame({"policy_tier": ["gold", "bronze", "silver"]})).ravel())
# [2. 0. 1.]

One exception to the whole rule. Tree-based models split on thresholds, so an ordinal encoding costs them far less than it costs a linear model. A tree can isolate category 2 with two splits. It is clumsy rather than wrong, and with high cardinality it is often the pragmatic choice.

High cardinality and categories you have never seen

Sixty vehicle makes is where one-hot encoding starts to hurt.

ohe_all = OneHotEncoder(handle_unknown="ignore", sparse_output=False).fit(cats)
print(ohe_all.transform(cats).shape)      # (1500, 67)

grouped = OneHotEncoder(handle_unknown="ignore", min_frequency=30,
                        sparse_output=False).fit(cats[["vehicle_make"]])
print(grouped.transform(cats[["vehicle_make"]]).shape)   # (1500, 12)
print(grouped.get_feature_names_out()[-1])
# vehicle_make_infrequent_sklearn

Sixty-seven columns from 1500 rows, where most of the new columns are nearly all zeros. min_frequency=30 collapses every make appearing fewer than 30 times into a single infrequent_sklearn column, cutting 60 categories to 12. That is usually a better dataset than the full expansion, because a category seen eleven times cannot support a reliable coefficient anyway.

Target encoding, which replaces each category with the average target value for that group, is the other standard answer at high cardinality. scikit-learn provides TargetEncoder with internal cross-fitting, and it must be used inside a pipeline, because computing group means on the full dataset puts the answer into the feature.

The unseen category problem

A new vehicle make will appear in production. What the encoder does then is set by one argument.

train_regions = pd.DataFrame({"region": ["North", "South", "East"]})
test_regions = pd.DataFrame({"region": ["North", "Central"]})

enc = OneHotEncoder(handle_unknown="ignore", sparse_output=False).fit(train_regions)
print(enc.get_feature_names_out())     # ['region_East' 'region_North' 'region_South']
print(enc.transform(test_regions))
# [[0. 1. 0.]
#  [0. 0. 0.]]

The unseen “Central” becomes a row of zeros, which means “none of the known categories” and lets the prediction proceed. Without handle_unknown="ignore" the same call raises a ValueError, and a ValueError in a scoring service at 2am is a worse outcome than a slightly uninformed prediction.

Set handle_unknown="ignore" on every one-hot encoder you deploy. For OrdinalEncoder the equivalent is handle_unknown="use_encoded_value" with an unknown_value you choose, such as -1.

Situation Encoder Key argument
Unordered, few categories OneHotEncoder handle_unknown="ignore"
Genuine order OrdinalEncoder categories=[[...]] in your order
Many categories, tree model OrdinalEncoder handle_unknown="use_encoded_value"
Many categories, linear model OneHotEncoder min_frequency or max_categories
Very high cardinality TargetEncoder Inside a pipeline, never fitted alone

A practical ordering for a new dataset: count the distinct values in every text column first, one-hot anything under about fifteen levels, give an explicit categories list to anything genuinely ordered, and only then decide what to do with the remainder. Most datasets have one or two awkward high-cardinality columns and a handful of easy ones, so the hard thinking stays confined to the few columns that deserve it.

Fit every encoder on training data only. An encoder learns which categories exist, and learning that from the full dataset is the same category of mistake as imputing across the split. Full parameter details are in the OneHotEncoder documentation.

Frequently Asked Questions

What is the difference between one-hot and label encoding?

One-hot creates a separate binary column per category, so no category is numerically larger than another. Label or ordinal encoding assigns integers, which implies order and distance. Use one-hot for unordered categories and ordinal only when the ranking is real, such as bronze, silver, gold.

How do you encode a column with hundreds of categories?

One-hot encoding becomes unaffordable. Group rare levels with min_frequency, or use TargetEncoder inside a pipeline so group means are computed with cross-fitting. For tree-based models, ordinal encoding is clumsy but workable and costs far less than it does for a linear model.

What happens if a new category appears after training?

By default scikit-learn raises a ValueError and your prediction fails. Set handle_unknown="ignore" on OneHotEncoder so unseen values become a row of zeros, or handle_unknown="use_encoded_value" with an unknown_value on OrdinalEncoder. New categories are certain in production.

Key Takeaways

  • Default to OneHotEncoder for unordered categories, because integer codes assert an order and a distance that a linear model will act on.
  • Pass your own categories list when an order genuinely exists, since the encoder otherwise sorts alphabetically and bronze, gold, silver is not the ranking you meant.
  • Set handle_unknown="ignore" before deploying anything, as an unseen category raises an error by default and new values are inevitable.
  • Group rare levels with min_frequency rather than expanding every category, which cut 60 vehicle makes to 12 columns in the example here.
  • Fit encoders on training folds only, because the set of known categories is itself something learned from data.