A model you can read out loud

Decision trees learn a sequence of yes-or-no questions. Is the EMI ratio above 0.40? If yes, is income below 3.31 lakh? Follow the answers down to a leaf and that leaf’s majority class is the prediction. No arithmetic, no coefficients, no scaling.

That readability is the main reason they remain in use despite being beaten on accuracy by almost everything in Section 8.

import numpy as np
import pandas as pd
from sklearn.tree import DecisionTreeClassifier, export_text
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(31)
n = 1200
X = pd.DataFrame({
    "income_lakh": rng.gamma(3, 2.2, n).clip(1, 40),
    "emi_ratio": rng.beta(2, 4, n),
    "credit_age_yrs": rng.integers(0, 25, n).astype(float),
})
z = -1.1 - 0.16 * X["income_lakh"] + 6.4 * X["emi_ratio"] - 0.09 * X["credit_age_yrs"]
y = pd.Series((rng.random(n) < 1 / (1 + np.exp(-z))).astype(int))
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=.3, random_state=0, stratify=y)

tree = DecisionTreeClassifier(max_depth=3, random_state=0).fit(X_tr, y_tr)
print(export_text(tree, feature_names=list(X.columns), max_depth=1))
# |--- emi_ratio <= 0.40
# |   |--- emi_ratio <= 0.21
# |   |--- emi_ratio >  0.21
# |--- emi_ratio >  0.40
# |   |--- income_lakh <= 3.31
# |   |--- income_lakh >  3.31

export_text prints the whole model. You can hand that to a credit officer and have a conversation about whether the thresholds are sensible, which is not possible with a set of coefficients and impossible with a neural network.

How decision trees choose their splits

At each node the algorithm tries every threshold on every feature and keeps the one that produces the purest two groups. Purity is measured by Gini impurity, which is the probability of misclassifying a randomly chosen row if you labelled it according to the group’s class proportions. A group that is entirely one class scores 0. A fifty-fifty split scores 0.5.

The alternative is entropy, which measures the same idea in units of information. In practice they rarely disagree.

for criterion in ["gini", "entropy", "log_loss"]:
    t = DecisionTreeClassifier(criterion=criterion, max_depth=4,
                               random_state=0).fit(X_tr, y_tr)
    print(f"{criterion:<9} test {t.score(X_te, y_te):.4f}")
# gini      test 0.7278
# entropy   test 0.7361
# log_loss  test 0.7361

A gap of 0.008, which is noise on 360 test rows. Do not spend time on this parameter.

The search is greedy: each split is chosen to be locally best with no consideration of what comes after. That is why a tree can miss a structure that two coordinated splits would capture easily, and why the result is not guaranteed to be the best possible tree.

Two properties follow from splitting on thresholds rather than distances. Scaling is irrelevant:

from sklearn.preprocessing import StandardScaler
scaler = StandardScaler().fit(X_tr)
raw = DecisionTreeClassifier(max_depth=4, random_state=0).fit(X_tr, y_tr).score(X_te, y_te)
scaled = (DecisionTreeClassifier(max_depth=4, random_state=0)
          .fit(scaler.transform(X_tr), y_tr).score(scaler.transform(X_te), y_te))
print(round(raw, 4), round(scaled, 4))   # 0.7278 0.7278

Identical, because multiplying a column by a constant does not change the ordering its splits depend on. That makes trees the one family where feature scaling can be skipped without argument. Extreme values are similarly harmless, which is why the outlier treatment in detecting and treating outliers mattered for linear models and not for trees.

Depth is the parameter that matters

An unconstrained tree keeps splitting until every leaf is pure, which means memorising the training set.

for d in [1, 2, 3, 5, 10, None]:
    t = DecisionTreeClassifier(max_depth=d, random_state=0).fit(X_tr, y_tr)
    print(f"depth {str(d):<5} train {t.score(X_tr, y_tr):.3f}"
          f"  test {t.score(X_te, y_te):.3f}  leaves {t.get_n_leaves()}")
# depth 1     train 0.727  test 0.725  leaves 2
# depth 2     train 0.754  test 0.744  leaves 4
# depth 3     train 0.777  test 0.756  leaves 8
# depth 5     train 0.823  test 0.722  leaves 28
# depth 10    train 0.950  test 0.692  leaves 140
# depth None  train 1.000  test 0.650  leaves 191

A perfect training score of 1.000 and a test score of 0.650. The unconstrained tree has 191 leaves for 840 training rows, so many leaves contain a handful of rows and describe noise. This is the clearest demonstration of overfitting in the course so far, and it matches the pattern set out in the main challenges of machine learning.

Depth 3 is best here, with eight leaves. Beyond that every additional split buys training accuracy and costs test accuracy.

max_depth is the blunt control. min_samples_leaf is often better, since it caps how small a group can be rather than how deep the tree goes, which lets the tree stay deep where data is plentiful. Setting min_samples_leaf=20 is a reasonable default on data of this size.

Pruning finds the limit for you

Cost complexity pruning grows a full tree and then removes the branches that justify their complexity least, controlled by ccp_alpha.

for a in [0.0, 0.001, 0.003, 0.006, 0.01, 0.02]:
    t = DecisionTreeClassifier(random_state=0, ccp_alpha=a).fit(X_tr, y_tr)
    print(f"ccp_alpha {a:<6} leaves {t.get_n_leaves():<4}"
          f" train {t.score(X_tr, y_tr):.3f} test {t.score(X_te, y_te):.3f}")
# ccp_alpha 0.0    leaves 191  train 1.000 test 0.650
# ccp_alpha 0.001  leaves 169  train 0.988 test 0.656
# ccp_alpha 0.003  leaves 17   train 0.833 test 0.742
# ccp_alpha 0.006  leaves 7    train 0.753 test 0.753
# ccp_alpha 0.01   leaves 3    train 0.754 test 0.744
# ccp_alpha 0.02   leaves 2    train 0.727 test 0.725

A clean curve with a peak at ccp_alpha=0.006 and seven leaves. Pruning reached 0.753 against the 0.756 that guessing max_depth=3 found, so it did not win here, but it got there by searching rather than by luck. Use cost_complexity_pruning_path to generate candidate values and cross-validate across them.

The weakness that motivates Section 8

for seed in [0, 1, 2]:
    idx = np.random.default_rng(seed).choice(len(X_tr), len(X_tr), replace=True)
    t = DecisionTreeClassifier(max_depth=3, random_state=0).fit(X_tr.iloc[idx], y_tr.iloc[idx])
    print(export_text(t, feature_names=list(X.columns), max_depth=0).split("n")[0])
# |--- emi_ratio <= 0.40
# |--- emi_ratio <= 0.40
# |--- emi_ratio <= 0.29

Three bootstrap samples of the same data, and the root threshold moved from 0.40 to 0.29. Here the chosen feature stayed constant, which is the mild version. With several correlated features of similar strength, the root feature itself changes between resamples, and the whole structure below it changes with it.

That instability is the defining weakness of a single tree. It is also the thing ensembles exploit: averaging many unstable trees produces a stable model, which is what Section 8 builds. A single tree is worth using when a human must read and approve the logic. When accuracy is the goal, it is a building block rather than a finished product. The splitting criteria are documented in the scikit-learn decision tree guide.

Frequently Asked Questions

How do decision trees decide where to split?

At every node the algorithm tests all thresholds on all features and picks the one producing the purest child groups, measured by Gini impurity or entropy. The search is greedy, so each split is locally best with no regard for later splits, which means the final tree is not guaranteed optimal.

What is the difference between gini and entropy?

Both measure how mixed the classes are within a group, with zero meaning pure. Gini is slightly faster because it avoids logarithms. They rarely disagree in practice: on the example above, the two differed by 0.008 on test accuracy, which is noise. Tune depth instead.

How do you stop a decision tree from overfitting?

Limit its growth. An unconstrained tree reached 1.000 training accuracy and 0.650 on test data with 191 leaves. Set max_depth or, better, min_samples_leaf so no leaf describes a handful of rows, and cross-validate ccp_alpha for cost complexity pruning.

Key Takeaways

  • Print the model with export_text and review the thresholds with someone who knows the domain, since readability is the main reason to prefer a single tree.
  • Constrain growth with min_samples_leaf or max_depth, because an unconstrained tree memorised the training set perfectly and lost a tenth of its test accuracy.
  • Skip feature scaling and outlier treatment entirely, as splits depend on ordering rather than distance.
  • Ignore the criterion parameter and spend the time on depth, given that gini and entropy differed by less than noise here.
  • Expect thresholds to move between resamples, and treat a single tree as a component for ensembles whenever accuracy matters more than readability.