One environment per project, always
Setting up Python for machine learning is a twenty-minute job that people turn into a recurring week-long problem by skipping one step: the virtual environment. Install packages globally and your third project will break your first, because pandas 3 changed behaviour that your January notebook relied on.
The whole setup is three commands. Create an isolated environment, activate it, install five packages into it.
python -m venv .venv
source .venv/bin/activate # Windows: .venvScriptsactivate
pip install -U scikit-learn pandas numpy matplotlib jupyterlab
That is it. scipy arrives as a scikit-learn dependency, so you rarely install it by name. Everything else in this chapter is about keeping that environment honest over the months you will use it.
Conda is the alternative, and worth choosing if you work on Windows, need non-Python dependencies such as a specific BLAS build, or expect to install packages that compile awkwardly. For most people on a recent Python, venv plus pip is simpler and faster.
What setting up Python for machine learning actually needs
Five libraries cover the entire course. Knowing which job belongs to which one saves you searching the wrong documentation.
| Package | What you use it for | Where it appears |
|---|---|---|
numpy |
Arrays, vectorised arithmetic, random number generation | Underneath everything |
pandas |
Loading, joining and reshaping tabular data | Data preparation |
scikit-learn |
Estimators, preprocessing, model selection, metrics | Modelling |
matplotlib |
Plots for exploration and diagnostics | Exploratory analysis |
jupyterlab |
Interactive notebooks | Exploration, not production |
Resist the urge to add more at the start. XGBoost, LightGBM, SHAP and PyTorch all have their place later in this course, and each one you install early is another version conflict waiting for you. Install a package when a chapter needs it.
One pairing to understand now, since it causes confusion: numpy provides the array type and scikit-learn consumes it, while pandas wraps arrays with labels. Most scikit-learn estimators accept either, and the array and matrix operations behind them are the same in both cases.
Verify before you trust it
Run this first, in the environment you just made. It takes a second and it tells you whether you are where you think you are.
import sys
import sklearn, numpy, pandas, scipy
print(sys.version.split()[0])
print("scikit-learn", sklearn.__version__)
print("numpy ", numpy.__version__)
print("pandas ", pandas.__version__)
print("scipy ", scipy.__version__)
# True means you are inside a virtual environment, False means you are not
print(sys.prefix != sys.base_prefix)
# 3.12.3
# scikit-learn 1.8.0
# numpy 2.4.4
# pandas 3.0.2
# scipy 1.17.1
# False
That last line is the one worth keeping. It printed False in the session above, meaning those packages went into the system Python rather than an isolated environment. If you see False after activating a venv, your editor or notebook kernel is pointed at a different interpreter, which is the single most common reason a package you definitely installed cannot be imported.
In Jupyter the fix is to select the kernel belonging to your environment, or to register it explicitly with python -m ipykernel install --user --name myproject.
Pin versions so your results survive
Reproducibility has two halves, and most people do only one of them.
Pin the packages. Freeze exactly what you have into a file that travels with the project.
pip freeze > requirements.txt
Anyone, including you in six months, recreates the environment with pip install -r requirements.txt. Without it, a rerun on a newer scikit-learn can produce different numbers and you will have no way to tell whether the model changed or the library did.
Seed the randomness. Many estimators are stochastic, and unseeded runs differ every time.
import numpy as np
from sklearn.datasets import load_wine
from sklearn.ensemble import RandomForestClassifier
data = load_wine()
a = RandomForestClassifier(n_estimators=50, random_state=0).fit(data.data, data.target)
b = RandomForestClassifier(n_estimators=50, random_state=0).fit(data.data, data.target)
c = RandomForestClassifier(n_estimators=50).fit(data.data, data.target)
print(np.array_equal(a.predict(data.data), b.predict(data.data))) # True
print(round(float(a.feature_importances_[0]), 6)) # 0.11067
print(round(float(c.feature_importances_[0]), 6)) # 0.123806
Same seed, identical model. No seed, and the importance of the first feature moves from 0.111 to 0.124 between runs. That gap is big enough to change which features you decide to keep, which is why random_state belongs on every estimator and every split you write. Treat an unseeded result as a result you cannot defend.
When the install goes wrong
Three failures account for most of it, and each has a one-line diagnosis.
A wheel fails to build and the terminal fills with compiler errors. You are on a Python release too new for the package, which has no prebuilt binary yet. Drop back one minor version.
pip installs successfully but the import still fails. Your shell pip and your notebook kernel belong to different environments. Run import sys; print(sys.executable) inside the notebook and compare it with which pip in the terminal.
Everything imports but results differ from a colleague’s. Compare sklearn.show_versions() output, which prints the library version alongside its numpy, scipy and BLAS details. Default parameters change between scikit-learn releases more often than people expect, and a changed default will move your numbers without changing a line of your code.
A layout that stays usable
churn-model/
.venv/ # never committed
data/raw/ # never committed, never edited
data/processed/
notebooks/ # exploration
src/ # functions you reuse
requirements.txt
README.md
Two rules make this worth the trouble. Raw data is read-only, so you can always rebuild from source. And anything a notebook does twice moves into src/, because code buried in cell 43 of an untitled notebook is code nobody will find again, including you. Notebooks are for exploration; the artefacts each project stage must hand over belong in files.
Commit requirements.txt and never commit data/raw/ or .venv/. Datasets in Git make the repository enormous, and raw data often carries the privacy obligations noted in collecting and sourcing data. Official platform-specific instructions are on the scikit-learn installation page.
Frequently Asked Questions
Which Python version should I use for machine learning?
Pick a stable release one or two versions behind the newest. The latest release often lacks compiled wheels for scientific packages, which turns a one-minute install into a compilation error. Check the scikit-learn installation page for the currently supported range before choosing.
Do I need Anaconda for machine learning?
No. Anaconda bundles the scientific stack and manages non-Python dependencies, which helps on Windows and in locked-down corporate environments. On macOS and Linux, python -m venv with pip installs the same packages faster and with less disk use. Either is fine; mixing them in one environment is not.
Why can Python not find a package I just installed?
Almost always because your interpreter and your installation target differ. You installed into a virtual environment but the notebook kernel points at system Python, or the reverse. Run print(sys.executable) to see exactly which interpreter is running, then install into that one.
Key Takeaways
- Create a virtual environment per project before installing anything, since a shared global environment guarantees a version conflict within a few months.
- Print
sys.prefix != sys.base_prefixafter setup to confirm you are inside the environment you think you are, which resolves most import failures immediately. - Commit a
requirements.txtproduced bypip freeze, or a future rerun on newer libraries will give different numbers with no way to attribute the change. - Set
random_stateon every estimator and split, because unseeded feature importances move enough between runs to change your decisions. - Keep raw data read-only and out of version control, and move any notebook code you use twice into a module.