An assumption nobody believes

A naive bayes classifier applies Bayes’ theorem to work out which class makes the observed features most probable. The word “naive” refers to the shortcut it takes: it assumes every feature is independent of every other, given the class.

That assumption is almost always false. Income and loan amount move together. The words “free” and “prize” appear in the same emails. The algorithm proceeds as though they do not, multiplying individual probabilities as if each feature arrived with no knowledge of the others.

It works anyway, much of the time, and understanding why is more useful than memorising the formula. Classification only requires the model to rank the classes correctly, not to estimate their probabilities accurately. The independence shortcut distorts the probabilities badly while often leaving the ranking intact.

The groundwork is in probability theory and distributions, where the base-rate effect and the posterior calculation are set out. Here the subject is which variant to use, what the assumption costs when it fails, and why the probabilities should not be trusted.

Which naive bayes classifier to use

Variant Feature type Typical use
GaussianNB Continuous Numeric tabular data
MultinomialNB Counts Word counts in text
BernoulliNB Binary Word present or absent
ComplementNB Counts Text with imbalanced classes

Choosing wrongly is a common and silent error. MultinomialNB on standardised numeric data will fail outright, because it cannot accept negative values. GaussianNB on word counts fits a bell curve to something that is mostly zeros.

Start with the loan data:

import numpy as np
import pandas as pd
from sklearn.naive_bayes import GaussianNB
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
from sklearn.pipeline import make_pipeline
from sklearn.model_selection import train_test_split
from sklearn.metrics import roc_auc_score

rng = np.random.default_rng(31)
n = 1200
X = pd.DataFrame({
    "income_lakh": rng.gamma(3, 2.2, n).clip(1, 40),
    "emi_ratio": rng.beta(2, 4, n),
    "credit_age_yrs": rng.integers(0, 25, n).astype(float),
})
z = -1.1 - 0.16 * X["income_lakh"] + 6.4 * X["emi_ratio"] - 0.09 * X["credit_age_yrs"]
y = pd.Series((rng.random(n) < 1 / (1 + np.exp(-z))).astype(int))

X_tr, X_te, y_tr, y_te = train_test_split(X, y, test_size=.3, random_state=0, stratify=y)

nb = GaussianNB().fit(X_tr, y_tr)
lr = make_pipeline(StandardScaler(), LogisticRegression(max_iter=5000)).fit(X_tr, y_tr)
print(round(roc_auc_score(y_te, nb.predict_proba(X_te)[:, 1]), 4))   # 0.8104
print(round(roc_auc_score(y_te, lr.predict_proba(X_te)[:, 1]), 4))   # 0.8059

The naive model narrowly beats logistic regression. It also trained in a single pass over the data, computing a mean and variance per feature per class, with no iteration at all. That speed is its main practical attraction.

What happens when the assumption breaks

Independence failing is not an abstract concern. Here is the cost, measured.

X_dup = X.copy()
for i in range(6):
    X_dup[f"emi_copy_{i}"] = X["emi_ratio"] + rng.normal(0, 0.001, n)   # near-duplicates

Xd_tr, Xd_te, yd_tr, yd_te = train_test_split(X_dup, y, test_size=.3,
                                              random_state=0, stratify=y)
nb2 = GaussianNB().fit(Xd_tr, yd_tr)
lr2 = make_pipeline(StandardScaler(),
                    LogisticRegression(max_iter=5000)).fit(Xd_tr, yd_tr)
print(round(roc_auc_score(yd_te, nb2.predict_proba(Xd_te)[:, 1]), 4))   # 0.7641
print(round(roc_auc_score(yd_te, lr2.predict_proba(Xd_te)[:, 1]), 4))   # 0.8058

Seven copies of the same information dropped naive bayes from 0.810 to 0.764. Logistic regression did not move. The model counted one piece of evidence seven times, because it has no way to notice that the columns agree.

That is the failure mode to watch for. Correlated features do not merely dilute a naive bayes classifier; they let one signal dominate in proportion to how many times it appears in your feature set. Before using it, drop near-duplicate columns.

The probabilities are not calibrated

p_nb = nb2.predict_proba(Xd_te)[:, 1]
p_lr = lr2.predict_proba(Xd_te)[:, 1]
print(round(float(((p_nb > 0.99) | (p_nb < 0.01)).mean()), 3))   # 0.678
print(round(float(((p_lr > 0.99) | (p_lr < 0.01)).mean()), 3))   # 0.006

Two thirds of the naive bayes predictions are above 0.99 or below 0.01. Logistic regression puts 0.6% of its predictions in those extremes on the same data.

Multiplying many probabilities together drives the result towards 0 or 1, and correlated features accelerate it. The ranking survives, which is why AUC stays respectable, but the numbers themselves are close to meaningless. Never use a raw naive bayes probability in an expected-value calculation or a capacity threshold without calibrating it first, and run the bucket check from reading a model’s output to see how far off it is.

Where it genuinely shines

Text classification, which is where this family earned its reputation.

from sklearn.feature_extraction.text import CountVectorizer
from sklearn.naive_bayes import MultinomialNB

subjects = ["win free cash now", "claim your free prize", "free money click here",
            "urgent you won a prize", "congratulations free gift card",
            "winner claim cash", "team standup at ten", "invoice for november",
            "can you review the draft", "lunch tomorrow at one",
            "notes from the client call", "reminder dentist friday"]
labels = [1] * 6 + [0] * 6

vec = CountVectorizer()
model = MultinomialNB().fit(vec.fit_transform(subjects), labels)

test = ["free cash for you", "review the november invoice", "claim your gift card"]
print(model.predict(vec.transform(test)))                      # [1 0 1]
print(model.predict_proba(vec.transform(test))[:, 1].round(3)) # [0.889 0.043 0.962]

Twelve training examples and it separates all three test cases correctly. The same problem appeared in machine learning vs traditional programming as the alternative to an unmaintainable rule file.

Text suits this algorithm for specific reasons. Vocabularies run to tens of thousands of columns, where the single-pass training cost matters and where methods that measure distance break down. Word counts are sparse, so most features are absent and contribute nothing. And while words genuinely are correlated, the correlation is spread thinly across a huge vocabulary rather than concentrated in a few duplicated columns.

Use it as a fast baseline on any text problem, as the first model when you have thousands of columns and little time, and when training must happen repeatedly on streaming data. Skip it when you need trustworthy probabilities, or when a handful of your features are strongly correlated. One practical note: encode categorical columns as described in encoding categorical variables before using the multinomial or Bernoulli variants. Full details are in the scikit-learn naive Bayes guide.

Frequently Asked Questions

Why is naive bayes called naive?

Because it assumes every feature is conditionally independent of every other given the class, which is nearly always false in real data. The name acknowledges the shortcut. Classification only needs the correct ranking of classes, so the assumption often damages the probabilities without changing the decision.

When should you use naive bayes instead of logistic regression?

For text and other very high-dimensional sparse data, when you need a baseline in seconds, or when training must repeat continuously. Choose logistic regression when features are correlated or when the probability values themselves drive a decision, since naive bayes probabilities are badly calibrated.

Which naive bayes variant should you use?

GaussianNB for continuous numeric features, MultinomialNB for word or event counts, BernoulliNB for binary presence indicators, and ComplementNB for text with imbalanced classes. Using the wrong one fails quietly, such as fitting a bell curve to count data that is mostly zeros.

Key Takeaways

  • Drop near-duplicate columns before fitting, because seven copies of one feature cut AUC from 0.810 to 0.764 while leaving logistic regression untouched.
  • Match the variant to your feature type, since MultinomialNB cannot accept negative values and GaussianNB fits poorly to sparse counts.
  • Treat the predicted probabilities as rankings rather than estimates, as two thirds of them landed above 0.99 or below 0.01 here.
  • Reach for it first on text classification, where a vocabulary of thousands of sparse columns plays to its single-pass training.
  • Use it as a baseline to beat rather than a final model on tabular data with a modest number of correlated features.