One environment per project, always

Setting up Python for machine learning is a twenty-minute job that people turn into a recurring week-long problem by skipping one step: the virtual environment. Install packages globally and your third project will break your first, because pandas 3 changed behaviour that your January notebook relied on.

The whole setup is three commands. Create an isolated environment, activate it, install five packages into it.

python -m venv .venv
source .venv/bin/activate        # Windows: .venvScriptsactivate
pip install -U scikit-learn pandas numpy matplotlib jupyterlab

That is it. scipy arrives as a scikit-learn dependency, so you rarely install it by name. Everything else in this chapter is about keeping that environment honest over the months you will use it.

Conda is the alternative, and worth choosing if you work on Windows, need non-Python dependencies such as a specific BLAS build, or expect to install packages that compile awkwardly. For most people on a recent Python, venv plus pip is simpler and faster.

What setting up Python for machine learning actually needs

Five libraries cover the entire course. Knowing which job belongs to which one saves you searching the wrong documentation.

Package What you use it for Where it appears
numpy Arrays, vectorised arithmetic, random number generation Underneath everything
pandas Loading, joining and reshaping tabular data Data preparation
scikit-learn Estimators, preprocessing, model selection, metrics Modelling
matplotlib Plots for exploration and diagnostics Exploratory analysis
jupyterlab Interactive notebooks Exploration, not production

Resist the urge to add more at the start. XGBoost, LightGBM, SHAP and PyTorch all have their place later in this course, and each one you install early is another version conflict waiting for you. Install a package when a chapter needs it.

One pairing to understand now, since it causes confusion: numpy provides the array type and scikit-learn consumes it, while pandas wraps arrays with labels. Most scikit-learn estimators accept either, and the array and matrix operations behind them are the same in both cases.

Verify before you trust it

Run this first, in the environment you just made. It takes a second and it tells you whether you are where you think you are.

import sys
import sklearn, numpy, pandas, scipy

print(sys.version.split()[0])
print("scikit-learn", sklearn.__version__)
print("numpy       ", numpy.__version__)
print("pandas      ", pandas.__version__)
print("scipy       ", scipy.__version__)

# True means you are inside a virtual environment, False means you are not
print(sys.prefix != sys.base_prefix)

# 3.12.3
# scikit-learn 1.8.0
# numpy        2.4.4
# pandas       3.0.2
# scipy        1.17.1
# False

That last line is the one worth keeping. It printed False in the session above, meaning those packages went into the system Python rather than an isolated environment. If you see False after activating a venv, your editor or notebook kernel is pointed at a different interpreter, which is the single most common reason a package you definitely installed cannot be imported.

In Jupyter the fix is to select the kernel belonging to your environment, or to register it explicitly with python -m ipykernel install --user --name myproject.

Pin versions so your results survive

Reproducibility has two halves, and most people do only one of them.

Pin the packages. Freeze exactly what you have into a file that travels with the project.

pip freeze > requirements.txt

Anyone, including you in six months, recreates the environment with pip install -r requirements.txt. Without it, a rerun on a newer scikit-learn can produce different numbers and you will have no way to tell whether the model changed or the library did.

Seed the randomness. Many estimators are stochastic, and unseeded runs differ every time.

import numpy as np
from sklearn.datasets import load_wine
from sklearn.ensemble import RandomForestClassifier

data = load_wine()

a = RandomForestClassifier(n_estimators=50, random_state=0).fit(data.data, data.target)
b = RandomForestClassifier(n_estimators=50, random_state=0).fit(data.data, data.target)
c = RandomForestClassifier(n_estimators=50).fit(data.data, data.target)

print(np.array_equal(a.predict(data.data), b.predict(data.data)))   # True
print(round(float(a.feature_importances_[0]), 6))                   # 0.11067
print(round(float(c.feature_importances_[0]), 6))                   # 0.123806

Same seed, identical model. No seed, and the importance of the first feature moves from 0.111 to 0.124 between runs. That gap is big enough to change which features you decide to keep, which is why random_state belongs on every estimator and every split you write. Treat an unseeded result as a result you cannot defend.

When the install goes wrong

Three failures account for most of it, and each has a one-line diagnosis.

A wheel fails to build and the terminal fills with compiler errors. You are on a Python release too new for the package, which has no prebuilt binary yet. Drop back one minor version.

pip installs successfully but the import still fails. Your shell pip and your notebook kernel belong to different environments. Run import sys; print(sys.executable) inside the notebook and compare it with which pip in the terminal.

Everything imports but results differ from a colleague’s. Compare sklearn.show_versions() output, which prints the library version alongside its numpy, scipy and BLAS details. Default parameters change between scikit-learn releases more often than people expect, and a changed default will move your numbers without changing a line of your code.

A layout that stays usable

churn-model/
    .venv/                  # never committed
    data/raw/               # never committed, never edited
    data/processed/
    notebooks/              # exploration
    src/                    # functions you reuse
    requirements.txt
    README.md

Two rules make this worth the trouble. Raw data is read-only, so you can always rebuild from source. And anything a notebook does twice moves into src/, because code buried in cell 43 of an untitled notebook is code nobody will find again, including you. Notebooks are for exploration; the artefacts each project stage must hand over belong in files.

Commit requirements.txt and never commit data/raw/ or .venv/. Datasets in Git make the repository enormous, and raw data often carries the privacy obligations noted in collecting and sourcing data. Official platform-specific instructions are on the scikit-learn installation page.

Frequently Asked Questions

Which Python version should I use for machine learning?

Pick a stable release one or two versions behind the newest. The latest release often lacks compiled wheels for scientific packages, which turns a one-minute install into a compilation error. Check the scikit-learn installation page for the currently supported range before choosing.

Do I need Anaconda for machine learning?

No. Anaconda bundles the scientific stack and manages non-Python dependencies, which helps on Windows and in locked-down corporate environments. On macOS and Linux, python -m venv with pip installs the same packages faster and with less disk use. Either is fine; mixing them in one environment is not.

Why can Python not find a package I just installed?

Almost always because your interpreter and your installation target differ. You installed into a virtual environment but the notebook kernel points at system Python, or the reverse. Run print(sys.executable) to see exactly which interpreter is running, then install into that one.

Key Takeaways

  • Create a virtual environment per project before installing anything, since a shared global environment guarantees a version conflict within a few months.
  • Print sys.prefix != sys.base_prefix after setup to confirm you are inside the environment you think you are, which resolves most import failures immediately.
  • Commit a requirements.txt produced by pip freeze, or a future rerun on newer libraries will give different numbers with no way to attribute the change.
  • Set random_state on every estimator and split, because unseeded feature importances move enough between runs to change your decisions.
  • Keep raw data read-only and out of version control, and move any notebook code you use twice into a module.