Most of the work is not modelling
Ask someone who has shipped three models where the time went and you will hear the same breakdown. Roughly half on getting data into a usable table, a quarter on deciding what to predict and whether anyone will act on it, and something like a tenth on choosing and tuning an algorithm. The machine learning project lifecycle is mostly plumbing and definition, with a short burst of modelling in the middle.
That ratio surprises people who learned from tutorials, where the data arrives clean and the task is already specified. It should shape how you plan. If you budget two weeks and assume eight days of modelling, you will be three weeks late.
The end-to-end predictive modelling workflow covers this ground from the business and statistical side. What follows is the engineering view: which artefact each stage hands to the next, and which stages send you backwards.
The stages of the machine learning project lifecycle
Seven stages, each producing something concrete that the next stage consumes.
| Stage | What it produces | Typical share of effort |
|---|---|---|
| Framing | A prediction task and a success threshold | 10% |
| Target definition | A precise, dated label per row | 5% |
| Data collection | A training table at one grain | 30% |
| Preparation | Cleaned features, encoded and scaled | 20% |
| Modelling | A fitted estimator and its hyperparameters | 10% |
| Evaluation | Held-out scores and an error analysis | 15% |
| Deployment and monitoring | A served model and drift alerts | 10% |
Those percentages vary by project, but the shape holds: the two stages everyone associates with machine learning, modelling and evaluation, account for about a quarter of the work between them.
The artefact column matters more than the percentages. A stage is finished when it has produced its artefact, not when it feels complete. “I have explored the data” is not a stage ending. “I have a table with one row per customer per month, 41 columns, 180,000 rows, covering January 2024 to March 2026” is.
What each handoff looks like
Framing hands modelling a sentence: predict whether a given customer will place an order in the next 90 days, and beat the current rule by at least 8 points of recall at the same contact volume. That sentence is what makes the task a supervised learning problem rather than an open question. Data collection hands preparation a table and a data dictionary. Modelling hands evaluation a fitted object plus the exact code that produced it. Deployment hands monitoring a model file, an input schema and a baseline distribution for every feature.
When a project goes wrong, the fault is almost always a handoff where the artefact was vague. Nobody wrote down the success threshold, so the model is “quite good” and nobody can say whether to ship it.
The loop in code
Stages four through six compress into one object. A Pipeline chains preprocessing and the estimator so that everything refits together on each fold, which is what stops test data influencing the preparation steps.
import numpy as np
import pandas as pd
from sklearn.pipeline import Pipeline
from sklearn.impute import SimpleImputer
from sklearn.preprocessing import StandardScaler
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import cross_val_score
rng = np.random.default_rng(4)
n = 900
loans = pd.DataFrame({
"months_employed": rng.integers(0, 180, n).astype(float),
"debt_to_income": rng.normal(0.34, 0.11, n),
"prior_arrears": rng.poisson(0.4, n).astype(float),
})
loans.loc[rng.choice(n, 70, replace=False), "months_employed"] = np.nan
risk = (-0.012 * loans["months_employed"].fillna(60)
+ 6.0 * loans["debt_to_income"]
+ 0.55 * loans["prior_arrears"])
loans["defaulted"] = (risk + rng.normal(0, 0.6, n) > 2.2).astype(int)
X = loans.drop(columns="defaulted")
y = loans["defaulted"]
pipe = Pipeline([
("impute", SimpleImputer(strategy="median")),
("scale", StandardScaler()),
("model", LogisticRegression(max_iter=1000)),
])
scores = cross_val_score(pipe, X, y, cv=5, scoring="roc_auc")
print(scores.round(3), round(scores.mean(), 3), round(scores.std(), 3))
# [0.897 0.904 0.946 0.923 0.948] 0.924 0.021
print(round(y.mean(), 3)) # 0.178
Twelve lines of modelling against a default rate of 17.8%. The spread across folds, 0.021, is worth as much attention as the mean: a model whose score varies by 0.05 between folds is telling you the dataset is too small to distinguish it from a worse model. Building pipelines properly has its own chapter later in this course; here the point is only that the middle of the lifecycle is small.
The loops nobody draws in the diagram
Lifecycle diagrams show arrows going one way. Real projects run three backward loops, and recognising which one you are in saves weeks.
The tight loop: modelling to preparation. Scores are mediocre, so you add a feature, change an encoding, retry. Minutes to hours per turn. This is the loop people enjoy and over-use, because it feels productive and rarely moves the number much after the first few turns.
The middle loop: evaluation to data collection. Error analysis shows the model fails on a specific segment, and the fix is more data or a different source rather than a better algorithm. Days per turn. This is usually the loop that actually improves things.
The outer loop: evaluation to framing. The model works and nobody can act on its output, or the metric improved and the business number did not. Weeks per turn, and it means the first stage was wrong. Expensive, which is why the framing stage deserves more than the 10% it usually gets.
A rule I would apply: before your third pass round the tight loop, force yourself to spend an hour on error analysis. Look at thirty wrong predictions individually. That hour routinely redirects the project to the middle or outer loop, where the real gain is.
Where projects stall
Three stalls account for most abandoned work, and each maps to a stage.
Data that does not exist at the grain you need, discovered in week three. A target nobody can define unambiguously, so two stakeholders mean different things by “churned”. A model that scores well and changes no decision, because the output arrives after the moment it would have been useful.
All three are findable in the first week by asking a single question at each stage: what exactly will this produce, and who consumes it? The failure modes in the main challenges of machine learning describe what goes wrong inside a model. These are what goes wrong around it.
Deployment deserves one warning. A model that lives in a notebook has not shipped, and the gap between a fitted estimator and a served prediction is larger than beginners expect: schema validation, latency budgets, retraining triggers, rollback. Section 11 covers it. Plan for it at the start rather than treating it as an afterthought, because deployment constraints sometimes rule out the model you were about to choose. Full options for chaining steps are in the Pipeline documentation.
Frequently Asked Questions
How long does a machine learning project take?
A first useful model on data that already exists typically takes four to eight weeks. Most of that is assembling and validating the training table, not modelling. Projects needing new data collection or labelling run considerably longer, since labelling cost scales with the number of examples required.
What are the stages of a machine learning project?
Framing the problem, defining the target, collecting data, preparing features, modelling, evaluating, then deploying and monitoring. Each stage produces a concrete artefact the next one consumes. Real projects loop backwards repeatedly, most usefully from evaluation back to data collection.
Why do most machine learning projects fail?
Rarely because of the algorithm. The common causes are data that does not exist at the required grain, a target definition nobody agreed on, and a prediction that arrives too late to change any decision. All three are detectable in the first week by asking what each stage will produce.
Key Takeaways
- Budget roughly half your schedule for data work and a tenth for modelling, because the reverse assumption is the most common cause of a late project.
- Define the concrete artefact each stage must hand over, since vague handoffs are where projects quietly go wrong.
- Spend an hour reading individual wrong predictions before your third round of model tweaking, as error analysis usually redirects you to a more valuable loop.
- Write down the success threshold during framing, or you will reach evaluation with no way to decide whether to ship.
- Check your deployment constraints before choosing a model, because latency and schema requirements can eliminate an estimator after you have tuned it.