Why you cannot just fit a line

Understanding how logistic regression works starts with seeing why the obvious alternative fails. Code loan default as 0 and 1, run linear regression on it, and the problem appears immediately.

import numpy as np
import pandas as pd
from sklearn.linear_model import LinearRegression
from sklearn.model_selection import train_test_split

rng = np.random.default_rng(31)
n = 1200
loans = pd.DataFrame({
    "income_lakh": rng.gamma(3, 2.2, n).clip(1, 40),
    "emi_ratio": rng.beta(2, 4, n),
    "credit_age_yrs": rng.integers(0, 25, n).astype(float),
})
z = (-1.1 - 0.16 * loans["income_lakh"] + 6.4 * loans["emi_ratio"]
     - 0.09 * loans["credit_age_yrs"])
loans["default"] = (rng.random(n) < 1 / (1 + np.exp(-z))).astype(int)

X, y = loans.drop(columns="default"), loans["default"]
X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=.3,
                                          random_state=0, stratify=y)

preds = LinearRegression().fit(X_tr, y_tr).predict(X_te)
print(round(float(preds.min()), 3), round(float(preds.max()), 3))   # -0.252 1.097
print(int(((preds < 0) | (preds > 1)).sum()), "of", len(preds))     # 38 of 360

Thirty-eight of 360 predictions fall outside the range a probability can occupy. A borrower with a minus 25% chance of defaulting is not a cautious estimate; it is a model producing nonsense. The fix is a single function applied to the output.

The statistical treatment of logistic regression covers the inferential side and the business framing. Here the subject is the mechanism: what the sigmoid does, what the coefficients mean afterwards, and which parameters you actually tune.

How logistic regression works: a linear score and a sigmoid

Logistic regression computes exactly the same weighted sum as linear regression, then passes it through the sigmoid function, which maps any real number into the interval between 0 and 1.

def sigmoid(v):
    return 1 / (1 + np.exp(-v))

for v in [-4, -2, 0, 2, 4]:
    print(f"sigmoid({v:>2}) = {sigmoid(v):.4f}")
# sigmoid(-4) = 0.0180
# sigmoid(-2) = 0.1192
# sigmoid( 0) = 0.5000
# sigmoid( 2) = 0.8808
# sigmoid( 4) = 0.9820

A score of zero becomes a probability of 0.5. Large positive scores approach 1 without reaching it, large negative scores approach 0. The curve is steepest near zero, which means the model is most sensitive to small changes in the score exactly where the decision is closest.

The quantity being modelled linearly is the log-odds. Odds are the probability of the event divided by the probability of it not happening, so a probability of 0.8 is odds of 4 to 1. Take the logarithm of that and you get a value that runs from minus infinity to plus infinity, which a linear combination of features can represent without hitting a boundary.

Training minimises log loss rather than squared error. Log loss charges you heavily for being confident and wrong: predicting 0.99 for a case that turns out negative costs far more than predicting 0.6 for it. There is no closed-form solution, so the fit runs iteratively using the kind of optimisation described in gradient descent explained.

from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
from sklearn.linear_model import LogisticRegression
from sklearn.metrics import roc_auc_score, log_loss

clf = make_pipeline(StandardScaler(),
                    LogisticRegression(max_iter=3000)).fit(X_tr, y_tr)
proba = clf.predict_proba(X_te)[:, 1]
print(round(roc_auc_score(y_te, proba), 4))   # 0.8059
print(round(log_loss(y_te, proba), 4))        # 0.4746

Reading the coefficients as odds

A logistic coefficient is a change in log-odds, which nobody can picture. Exponentiate it and you get an odds ratio, which multiplies the odds.

raw = LogisticRegression(max_iter=5000).fit(X_tr, y_tr)
coef = pd.Series(raw.coef_[0], index=X.columns)
print(coef.round(4).to_dict())
# {'income_lakh': -0.1456, 'emi_ratio': 5.0012, 'credit_age_yrs': -0.0687}
print(np.exp(coef).round(4).to_dict())
# {'income_lakh': 0.8645, 'emi_ratio': 148.5902, 'credit_age_yrs': 0.9336}

Each extra lakh of income multiplies the odds of default by 0.86, a 14% reduction. Each extra year of credit history multiplies them by 0.93.

The emi_ratio figure of 148.59 looks absurd until you notice the units. That column runs from 0 to 1, so the odds ratio describes going from no EMI burden to spending your entire income on repayments. A tenth of that range is exp(0.5), about 1.65, which is the number worth quoting. Always state the unit change an odds ratio refers to, and remember that comparing raw coefficients across columns requires standardising first, as reading a model’s output showed.

The C parameter

LogisticRegression applies an L2 penalty by default, and C controls it inversely: small C means heavy regularisation.

for C in [0.001, 0.01, 0.1, 1, 100]:
    m = make_pipeline(StandardScaler(),
                      LogisticRegression(C=C, max_iter=5000)).fit(X_tr, y_tr)
    print(f"C={C:<7} AUC {roc_auc_score(y_te, m.predict_proba(X_te)[:, 1]):.4f}"
          f"   sum|coef| {np.abs(m[-1].coef_[0]).sum():.3f}")
# C=0.001   AUC 0.8050   sum|coef| 0.249
# C=0.01    AUC 0.8054   sum|coef| 1.148
# C=0.1     AUC 0.8053   sum|coef| 2.045
# C=1       AUC 0.8059   sum|coef| 2.276
# C=100     AUC 0.8059   sum|coef| 2.308

AUC moves by 0.0009 across five orders of magnitude while the coefficients shrink by a factor of nine. With 840 training rows and three features there is nothing to overfit, so the penalty has nothing to do. Tuning C matters when features approach or exceed rows, and barely at all otherwise.

Because the penalty acts on raw coefficient size, scaling is required rather than optional here, not merely helpful. The default lbfgs solver handles most problems; liblinear suits small datasets and saga is the one that supports L1 penalties on large data. If the fit warns about convergence, raise max_iter before suspecting anything else, since the default of 100 is too low for most real feature sets.

Two limits to name. Logistic regression draws a straight boundary between classes, so a circular or XOR-shaped separation needs explicit interaction terms before it can be learned at all. Adding squared terms and pairwise products is usually enough, and if it is not, that is the signal to move to a tree-based or kernel method rather than to keep patching the feature set. And when one class is rare, the fitted probabilities are systematically too low, which resampling and class weights addresses and Section 6 revisits later. Parameter details are in the LogisticRegression reference.

Frequently Asked Questions

Why can you not use linear regression for classification?

It produces values outside the range 0 to 1, which cannot be read as probabilities. In the example above, 38 of 360 predictions fell outside that range, including negative ones. Linear regression also minimises squared error, which penalises confident correct predictions alongside confident wrong ones.

What does the sigmoid function do in logistic regression?

It maps any real number into the interval between 0 and 1, turning an unbounded linear score into a probability. A score of 0 becomes 0.5, a score of 2 becomes 0.88, and the curve is steepest near zero, where the model is closest to being undecided.

How do you interpret logistic regression coefficients?

Exponentiate them to get odds ratios. A coefficient of -0.1456 becomes 0.8645, meaning each extra unit multiplies the odds by 0.86. Always state the unit change the ratio refers to, and standardise your features before comparing coefficient sizes across columns.

Key Takeaways

  • Reach for logistic regression rather than linear regression on any binary target, since a linear fit produces impossible probabilities outside 0 to 1.
  • Exponentiate coefficients into odds ratios before explaining them, and always name the unit change so a value like 148 is not quoted out of context.
  • Scale your features before fitting, because the default L2 penalty acts on raw coefficient magnitudes.
  • Spend little time tuning C when rows comfortably exceed columns, as AUC moved by 0.0009 across five orders of magnitude here.
  • Add interaction or polynomial terms when classes are not separable by a straight boundary, since logistic regression draws only one.