Six lines, and then the interesting part

Your first machine learning model takes about six lines of scikit-learn. Load some data, split it, create an estimator, call .fit(), call .score(). The library is deliberately boring: every model in it works the same way, so learning one means learning all of them.

What makes this chapter worth reading is the part after those six lines. The first run scores 0.667. One preprocessing step takes it to 0.933, with the same algorithm and the same data, and understanding why that happened teaches you more than any amount of model selection.

The end-to-end predictive modelling workflow walks the same territory with a business and statistics emphasis. Here the subject is the scikit-learn API itself: what each method does, what it returns, and what the interface assumes about your data.

The contract every estimator shares

Three method names cover almost everything in the library.

Method Who has it What it does
.fit(X, y) Every estimator Learns parameters from training data
.predict(X) Models Returns one prediction per row
.score(X, y) Models Returns a default metric, accuracy or R squared
.transform(X) Preprocessors Returns modified data
.fit_transform(X) Preprocessors Fits and transforms in one call

The split between models and preprocessors matters. A StandardScaler has .fit() and .transform() but no .predict(), because it produces data rather than answers. A KNeighborsClassifier has .predict() but no .transform().

Two expectations about your inputs. X is two-dimensional, shaped as (n_samples, n_features), and y is one-dimensional with one label per row in matching order. This is the labelled structure supervised learning requires, and getting the order wrong raises no error at all.

Building your first machine learning model

The wine dataset ships with scikit-learn, so there is nothing to download. It has 178 rows, 13 chemical measurements, and three cultivars to tell apart.

from sklearn.datasets import load_wine
from sklearn.model_selection import train_test_split
from sklearn.neighbors import KNeighborsClassifier

wine = load_wine(as_frame=True)
X, y = wine.data, wine.target

print(X.shape, y.shape)                                 # (178, 13) (178,)
print(X.columns.tolist()[:5])
# ['alcohol', 'malic_acid', 'ash', 'alcalinity_of_ash', 'magnesium']
print(y.value_counts().sort_index().to_dict())          # {0: 59, 1: 71, 2: 48}

X_train, X_test, y_train, y_test = train_test_split(
    X, y, test_size=0.25, random_state=0, stratify=y)
print(X_train.shape, X_test.shape)                      # (133, 13) (45, 13)

model = KNeighborsClassifier(n_neighbors=5)
model.fit(X_train, y_train)
print(round(model.score(X_test, y_test), 3))            # 0.667

Two arguments in that split deserve attention now rather than later. random_state=0 makes the split reproducible. stratify=y keeps the three class proportions the same in both halves, which matters here because the smallest class has only 48 rows and an unlucky random split could leave very few of them in training. The reasoning behind splitting at all belongs to Section 7; for now, treat it as the rule that you never score a model on rows it has already seen.

So, 66.7%. Against three roughly balanced classes, random guessing would score about 33%, so the model has learned something. It is also clearly not good.

Why the score jumped to 0.933

KNeighborsClassifier predicts by finding the five most similar rows in the training set and taking a vote. Similar means close in straight-line distance across all 13 columns.

Look at the units in this dataset. proline runs into the hundreds. magnesium is around 100. hue sits below 2. When you compute a distance across those columns, proline contributes thousands and hue contributes fractions, so the model is effectively clustering on proline alone and ignoring most of the information you gave it.

from sklearn.preprocessing import StandardScaler

scaler = StandardScaler().fit(X_train)          # fit on training data only

model = KNeighborsClassifier(n_neighbors=5)
model.fit(scaler.transform(X_train), y_train)
print(round(model.score(scaler.transform(X_test), y_test), 3))   # 0.933

print(model.predict(scaler.transform(X_test))[:8])  # [2 0 1 0 2 2 2 2]
print(y_test.to_numpy()[:8])                        # [2 0 1 0 2 2 2 2]

From 0.667 to 0.933. No change of algorithm, no tuning, no extra data. Every column now contributes on comparable terms.

Notice scaler.fit(X_train) rather than scaler.fit(X). The scaler learns a mean and a standard deviation, and learning those from the full dataset would let test rows influence the training process. Feature scaling has its own chapter in Section 4; the habit to form today is that anything which learns from data gets fitted on the training split only.

This is also why Pipeline exists, chaining the scaler and the model into one object so the fit order cannot go wrong. Section 10 covers it properly.

What this run has and has not shown

It has shown that the API works and that preprocessing can matter more than model choice. It has not shown that you have a good model, for three reasons worth naming.

Forty-five test rows is a small test set, so a single swing of three predictions moves the score by seven points. The accuracy figure alone hides which classes are being confused with which. And n_neighbors=5 was a guess rather than a choice.

None of those are failures of this chapter. They are the next three chapters: how to read what the model is actually telling you, how to validate honestly, and how to select hyperparameters.

There is a habit worth starting now. Before accepting any score, rerun the whole block with random_state=1 and then random_state=2. If 0.933 becomes 0.867 and then 0.911, you have learned that your test set is too small to separate a good model from a slightly worse one, and that single piece of information will save you from arguing over differences that are not there. What matters here is that the target you fed in came from a decision you made during target definition, and the score you just printed is only meaningful against the baseline you set during problem framing.

Swapping the estimator is a one-line change, which is the point of the shared interface. Replace KNeighborsClassifier(n_neighbors=5) with LogisticRegression(max_iter=5000) or RandomForestClassifier(random_state=0) and everything else stays identical. Parameter details are in the references for KNeighborsClassifier and train_test_split.

Frequently Asked Questions

What is the difference between fit and predict in scikit-learn?

.fit(X, y) learns from labelled training data and stores the result inside the estimator. .predict(X) applies what was learned to new rows and returns one answer each. You call fit once per training run and predict as often as you have new data.

Why does scaling change my model’s accuracy so much?

Distance-based algorithms such as k-nearest neighbours measure similarity across all columns at once. A column measured in hundreds overwhelms one measured in decimals, so the model effectively ignores most features. Scaling puts every column on comparable terms. Tree-based models are unaffected.

Which model should I start with in scikit-learn?

Start with LogisticRegression for classification or LinearRegression for regression. Both are fast, give readable coefficients, and set an honest baseline. Move to RandomForestClassifier once you have that number. Starting with a complex model hides whether the extra complexity bought anything.

Key Takeaways

  • Learn the fit, predict, transform contract once, because every estimator in scikit-learn follows it and swapping models becomes a one-line change.
  • Pass stratify=y when splitting classification data, so a small class cannot end up badly represented in your training half.
  • Scale your features before any distance-based model, since unit differences alone moved accuracy from 0.667 to 0.933 on this dataset.
  • Fit scalers and other preprocessors on the training split only, never on the full dataset, or test information leaks into training.
  • Treat a single accuracy number on 45 test rows as a smoke test rather than evidence, and set random_state so the run is repeatable.