predict is the least informative thing a model gives you
.predict() returns a class label and throws away everything that made it interesting. Interpreting model output properly means reaching for the other methods: the probabilities behind the label, the coefficients that produced them, and the importances that say which columns mattered.
A fitted estimator is not a black box by default. A logistic regression will tell you exactly how much each feature shifted its decision, in units you can check. A random forest will rank its features. Both will give you the confidence behind every prediction rather than just the verdict, and that confidence is usually what the business actually needs.
Everything below continues from the wine classifier built in your first model in scikit-learn, restated here so this page runs on its own.
from sklearn.datasets import load_wine
from sklearn.model_selection import train_test_split
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
wine = load_wine(as_frame=True)
X, y = wine.data, wine.target
X_train, X_test, y_train, y_test = train_test_split(
X, y, test_size=0.25, random_state=0, stratify=y)
scaler = StandardScaler().fit(X_train)
clf = LogisticRegression(max_iter=5000).fit(scaler.transform(X_train), y_train)
X_test_s = scaler.transform(X_test)
The three ways a classifier answers
print(clf.classes_) # [0 1 2]
print(clf.predict(X_test_s)[:5]) # [2 0 1 0 2]
proba = clf.predict_proba(X_test_s)
print(proba.shape) # (45, 3)
print(proba[:3].round(3))
# [[0.001 0.001 0.997]
# [0.867 0.132 0.001]
# [0.001 0.998 0.002]]
print(clf.decision_function(X_test_s)[:2].round(3))
# [[-2.117 -2.292 4.409]
# [ 2.842 0.96 -3.801]]
predict_proba returns one column per class, in the order given by classes_. That ordering is the detail that catches people: scikit-learn sorts classes, so for labels 0 and 1 the positive class is column index 1, and for string labels "churn" and "stay" it is alphabetical. Never hard-code [:, 1] without checking classes_ first.
Each row sums to 1. The second test row is 0.867 for class 0 and 0.132 for class 1, which is a genuinely uncertain prediction that .predict() reports simply as 0. Those two numbers are the ones a reviewer wants.
decision_function returns the raw scores before they are squashed into probabilities. You need it for support vector machines, which have no predict_proba by default, and it is sometimes more stable for ranking. For ordinary work, probabilities are easier to explain to anyone else.
Interpreting model output from coefficients
For linear models, coef_ tells you the direction and size of each feature’s effect. There is a trap in the shape.
print(clf.coef_.shape) # (3, 13)
print(clf.intercept_.round(3)) # [ 0.368 0.744 -1.112]
Three rows, not one. A multi-class logistic regression fits one set of coefficients per class, so clf.coef_[0] describes what pushes a sample towards class 0 specifically. For binary classification you get a single row and coef_[0] is the whole model.
Now the part that changes how you read these numbers. Coefficients are in the units of their feature, so comparing them is meaningless unless the features share a scale.
import pandas as pd
raw = LogisticRegression(max_iter=20000).fit(X_train, y_train) # unscaled
scaled = LogisticRegression(max_iter=5000).fit(scaler.transform(X_train), y_train)
comp = pd.DataFrame({"feature": X.columns,
"unscaled": raw.coef_[0].round(4),
"scaled": scaled.coef_[0].round(3)})
print(comp.reindex(comp["scaled"].abs().sort_values(ascending=False).index)
.head(4).to_string(index=False))
# feature unscaled scaled
# proline 0.0087 1.032
# alcalinity_of_ash -0.2636 -0.808
# alcohol 0.4292 0.748
# flavanoids 0.7812 0.694
Read the proline row. Its unscaled coefficient is 0.0087, small enough that you would dismiss it. On standardised features it is 1.032, the largest in the model. The coefficient was tiny only because proline is measured in hundreds while the other columns are not.
The rule that follows: standardise before comparing coefficients, or you are ranking your features by their units of measurement. A coefficient’s sign is reliable either way; its magnitude is not.
Feature importances, and what they do not mean
Tree-based models expose feature_importances_ instead.
from sklearn.ensemble import RandomForestClassifier
rf = RandomForestClassifier(n_estimators=300, random_state=0).fit(X_train, y_train)
imp = pd.Series(rf.feature_importances_, index=X.columns).sort_values(ascending=False)
print(imp.head(5).round(4).to_dict())
# {'proline': 0.1959, 'color_intensity': 0.1709, 'flavanoids': 0.1654,
# 'od280/od315_of_diluted_wines': 0.1095, 'alcohol': 0.1089}
print(round(float(imp.sum()), 4)) # 1.0
The forest independently ranks proline first, agreeing with the standardised coefficients. Agreement between two different model families is reasonable evidence that the signal is real.
Three limits to state plainly. These values always sum to 1, so they are shares of the model’s splitting activity rather than absolute effect sizes. They say nothing about direction, so a high importance does not tell you whether more proline means class 0 or class 2. And the default impurity-based calculation is biased towards high-cardinality features, which means a near-unique ID column will look important while contributing nothing generalisable. Permutation importance avoids that bias, and SHAP values give per-prediction attributions; both come later in this course.
Correlated features split their importance between them. Two columns carrying the same information may each score 0.05 while one alone would score 0.10, so a low importance does not prove a feature is useless.
Are the probabilities worth believing?
A probability of 0.8 should mean that roughly 80% of such cases turn out positive. Models do not guarantee this, so check rather than assume.
import numpy as np
from sklearn.datasets import load_breast_cancer
bc = load_breast_cancer(as_frame=True)
Xb_tr, Xb_te, yb_tr, yb_te = train_test_split(
bc.data, bc.target, test_size=0.4, random_state=1, stratify=bc.target)
rf2 = RandomForestClassifier(n_estimators=200, random_state=1).fit(Xb_tr, yb_tr)
p = rf2.predict_proba(Xb_te)[:, 1]
for lo, hi in [(0.0, 0.2), (0.2, 0.4), (0.4, 0.6), (0.6, 0.8), (0.8, 1.01)]:
mask = (p >= lo) & (p < hi)
if mask.sum():
print(f"{lo}-{hi}", int(mask.sum()),
round(float(p[mask].mean()), 3),
round(float(yb_te.to_numpy()[mask].mean()), 3))
# 0.0-0.2 71 0.036 0.0
# 0.2-0.4 9 0.298 0.333
# 0.4-0.6 7 0.486 0.429
# 0.6-0.8 8 0.715 0.75
# 0.8-1.01 133 0.98 0.985
Predicted against observed, bucket by bucket. Here they track closely: the 0.6 to 0.8 bucket predicted 0.715 and 0.75 of those cases were actually positive. This forest is well calibrated on this data, which is a finding rather than a given. Boosted trees and support vector machines frequently are not, and an uncalibrated 0.9 can mean anything.
Run this check whenever the probability itself drives a decision, such as expected value calculations or a threshold set by review capacity. If it fails, calibration methods exist to fix it. The probability groundwork explains why a predicted probability is a long-run frequency claim rather than a confidence score, and the capacity-based thresholds from problem framing depend on it holding. Method details are in the references for LogisticRegression and RandomForestClassifier.
Frequently Asked Questions
What is the difference between predict and predict_proba?
.predict() returns a single class label per row. .predict_proba() returns one probability per class, in the order shown by .classes_, and the row sums to 1. Probabilities let you set your own threshold and see which predictions were borderline, which the label alone hides.
How do you interpret coefficients in logistic regression?
The sign shows direction and the magnitude shows strength, but only if the features were standardised first. Unscaled coefficients reflect each feature’s units, so a variable measured in hundreds gets a tiny coefficient regardless of importance. Multi-class models produce one coefficient row per class.
Are feature importances reliable in random forests?
Directionally useful, with caveats. Default impurity-based importances favour high-cardinality features and split their value between correlated columns, so a low score does not prove a feature is useless. Permutation importance is the more trustworthy alternative, and importances never indicate direction.
Key Takeaways
- Check
.classes_before indexingpredict_probaoutput, because scikit-learn sorts labels and the positive class is not always the second column. - Standardise your features before comparing coefficient magnitudes, since on raw data you are ranking variables by their units rather than their effect.
- Expect one coefficient row per class from multi-class models, and read
coef_[k]as what pushes a sample towards classk. - Treat feature importances as shares of splitting activity with no direction, and distrust them when features are correlated or high-cardinality.
- Bucket your predicted probabilities against observed rates before letting any probability drive a decision, rather than assuming the model is calibrated.