A ratio beats a better algorithm

Two raw columns, amount and credit_limit, give a fraud model an AUC of 0.613. Divide one by the other, add the hour of day, and the same model reaches 0.909. Feature engineering did that, not a change of algorithm, not more data, and not tuning.

This is the highest-return activity in applied machine learning and it is almost entirely unglamorous. You are translating what a domain expert already knows into columns a model can read. A fraud investigator does not care that a transaction was 40,000 rupees; they care that it was 80% of the available limit at 3am.

Models cannot discover that for you. A linear model can only add up the columns it is given, and even a gradient boosting machine would need many splits to approximate a ratio it could have been handed directly.

The transformations that earn their keep

Five families cover most of what you will build.

Ratios and differences. Amount over limit, spend this month over spend last month, price over category median. Ratios normalise across customers of different sizes, which is why they work so often.

Datetime decomposition. A timestamp is useless to a model as a number. Hour of day, day of week, month, and flags such as weekend or holiday are where the signal lives.

Aggregations by group. How does this transaction compare with this customer’s usual behaviour? Group means, counts and standard deviations turn an isolated row into a contextual one.

Interactions. Two columns whose combination matters more than either alone. PolynomialFeatures generates these mechanically, though hand-picked ones are usually better.

Binning. Converting a continuous column into bands, which lets a linear model express a non-monotonic relationship. Risk bands by age are the standard example: the relationship between age and default is rarely a straight line, and three bands capture a shape a single coefficient cannot.

import numpy as np
import pandas as pd
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.pipeline import make_pipeline
from sklearn.model_selection import cross_val_score

rng = np.random.default_rng(33)
m = 4000
tx = pd.DataFrame({
    "customer_id": rng.integers(0, 500, m),
    "ts": pd.to_datetime("2026-01-01") + pd.to_timedelta(
        rng.integers(0, 120 * 24 * 60, m), unit="m"),
    "amount": rng.gamma(1.4, 3000, m),
    "credit_limit": rng.choice([50000, 100000, 200000, 500000], m),
}).sort_values("ts").reset_index(drop=True)

hour = tx["ts"].dt.hour
ratio = tx["amount"] / tx["credit_limit"]
tx["is_fraud"] = ((2.6 * ratio
                   + 0.9 * ((hour < 5) | (hour > 22)).astype(float)
                   + 0.4 * rng.normal(0, 1, m)) > 1.35).astype(int)
y = tx["is_fraud"]
print("fraud rate", round(float(y.mean()), 3))      # 0.056

def auc(frame, cols):
    pipe = make_pipeline(StandardScaler(), LogisticRegression(max_iter=3000))
    return round(cross_val_score(pipe, frame[cols], y, cv=5, scoring="roc_auc").mean(), 4)

print("raw columns:", auc(tx, ["amount", "credit_limit"]))   # 0.6128

Feature engineering, applied

eng = tx.copy()
eng["amount_to_limit"] = eng["amount"] / eng["credit_limit"]
eng["hour"] = eng["ts"].dt.hour
eng["is_night"] = ((eng["hour"] < 5) | (eng["hour"] > 22)).astype(int)
eng["dayofweek"] = eng["ts"].dt.dayofweek
eng["is_weekend"] = (eng["dayofweek"] >= 5).astype(int)
eng["amount_vs_cust_mean"] = (eng["amount"]
                              / eng.groupby("customer_id")["amount"].transform("mean"))

features = ["amount", "credit_limit", "amount_to_limit", "hour",
            "is_night", "dayofweek", "is_weekend", "amount_vs_cust_mean"]

print("all eight:", auc(eng, features))                              # 0.9089
print("is_night alone:", auc(eng, ["is_night"]))                     # 0.8699
print("ratio + night:", auc(eng, ["amount_to_limit", "is_night"]))   # 0.9126

Three numbers worth reading together. Eight engineered features give 0.909. One feature, a binary night flag, gives 0.870 on its own. And two well-chosen features give 0.913, beating all eight.

That last result is the lesson. Adding features is not the same as adding signal, and six of the eight columns contributed nothing except an opportunity to overfit. Feature engineering means finding the right transformations, not generating many of them.

Cyclical features need care

Hour of day as an integer tells the model that 23 and 0 are 23 units apart, when they are one hour apart. Encoding the hour as a position on a circle fixes it.

eng["hour_sin"] = np.sin(2 * np.pi * eng["hour"] / 24)
eng["hour_cos"] = np.cos(2 * np.pi * eng["hour"] / 24)

print(eng.loc[eng["hour"].isin([23, 0, 1]), ["hour", "hour_sin", "hour_cos"]]
      .drop_duplicates("hour").round(3).to_string(index=False))
# hour  hour_sin  hour_cos
#    0     0.000     1.000
#   23    -0.259     0.966
#    1     0.259     0.966

Hours 23, 0 and 1 now sit close together in the sine and cosine pair, which is the truth about a clock. The same trick applies to day of week and month. It matters for linear and distance-based models; a tree can isolate hours 23 and 0 with splits and largely does not need it.

Where feature creation goes wrong

Features that use the future. The most expensive mistake in this chapter. A customer-level average computed over the whole dataset includes transactions that happened after the row you are predicting. The aggregation above is a simplified example and would need a point-in-time computation in production, as described in collecting and sourcing data. Use expanding or rolling windows that only look backwards.

Features nobody can compute at prediction time. A beautiful column derived from a system that updates nightly is useless to a model scoring at checkout. Ask when each input becomes available before building anything on it.

Generating features mechanically. PolynomialFeatures(degree=2) on 30 columns produces 495 of them, mostly noise, and with high cardinality categorical columns encoded as in encoding categorical variables the count grows faster still. Automated interaction generation occasionally helps and usually buries the signal.

Target-derived features. Computing a feature from the target, including the mean target for a group, requires cross-fitting. Done naively, the answer ends up inside the input.

A workflow that works: ask the domain expert what they look at, build those columns by hand, measure each addition, and stop when a feature stops paying. The diagnostics from exploratory data analysis tell you which raw columns are worth transforming in the first place. Mechanical interaction generation is documented under PolynomialFeatures.

Frequently Asked Questions

What is feature engineering in machine learning?

Creating new input columns from existing ones so a model can use information it could not extract itself. Typical examples are ratios, differences, datetime parts and group aggregations. It usually improves results more than changing algorithms, as two derived columns here moved AUC from 0.613 to 0.913.

How do you create features from a date column?

Extract the parts that carry meaning: hour, day of week, month, and flags for weekend or holiday. A raw timestamp as a number is close to useless. For hour and month, also add sine and cosine encodings so the model knows 23:00 and 00:00 are adjacent rather than far apart.

Does feature engineering still matter with gradient boosting?

Yes, less than for linear models but far from zero. Boosting can approximate a ratio through many splits, which costs depth and data. Handing it the ratio directly is cheaper and generalises better. Domain features remain the main lever on tabular problems.

Key Takeaways

  • Build ratios and group comparisons first, because normalising a value against a limit or a customer’s own history is where most tabular signal comes from.
  • Decompose every timestamp into hour, day of week and weekend flags, and add sine and cosine encodings so cyclical values wrap correctly.
  • Measure each feature’s contribution rather than assuming more is better, since two chosen columns beat all eight in the example here.
  • Confirm that every input is knowable at prediction time and computed from past data only, or you have built a column that cannot exist in production.
  • Prefer a handful of features a domain expert would recognise over mechanically generated polynomial interactions.