Looking at data to decide what to fix
An analyst explores data to answer questions about the business. You are exploring it to answer one question about the data itself: what will break when I put this into a model? Exploratory data analysis for machine learning is a diagnosis rather than a tour, and it ends with a list of repairs, not a set of charts.
That changes what you look at. Average claim value by region is interesting to a business and irrelevant to you right now. Whether region has four categories or four thousand decides whether one-hot encoding is viable. Whether survey_score is missing at random or missing for a specific kind of customer decides whether imputing it will mislead the model.
Six things to establish, in order. After that the rest of Section 4 does the actual fixing, so resist treating anything yet.
What exploratory data analysis for machine learning has to establish
Start with a profile rather than a plot. One table, every column, three facts each.
import numpy as np
import pandas as pd
rng = np.random.default_rng(12)
n = 1500
df = pd.DataFrame({
"policy_id": [f"P{100000+i}" for i in range(n)],
"region": rng.choice(["North", "South", "East", "West"], n, p=[.4, .3, .2, .1]),
"vehicle_make": rng.choice([f"make_{i}" for i in range(60)], n),
"driver_age": rng.integers(18, 85, n).astype(float),
"annual_mileage": rng.gamma(2.0, 6000, n),
"prior_claims": rng.poisson(0.5, n).astype(float),
"policy_tier": rng.choice(["bronze", "silver", "gold"], n, p=[.5, .35, .15]),
"survey_score": rng.normal(7, 1.5, n),
"country": "IN",
})
df.loc[rng.choice(n, 400, replace=False), "survey_score"] = np.nan
older = df.index[df["driver_age"] > 65]
df.loc[rng.choice(older, int(len(older) * 0.55), replace=False), "annual_mileage"] = np.nan
risk = (0.02 * (80 - df["driver_age"]) + 0.00004 * df["annual_mileage"].fillna(12000)
+ 0.6 * df["prior_claims"])
df["claimed"] = (risk + rng.normal(0, 0.5, n) > 1.55).astype(int)
df = pd.concat([df, df.iloc[:25]], ignore_index=True) # 25 accidental duplicates
profile = pd.DataFrame({
"dtype": df.dtypes.astype(str),
"missing_pct": (df.isna().mean() * 100).round(1),
"n_unique": df.nunique(),
})
print(profile.to_string())
# dtype missing_pct n_unique
# policy_id str 0.0 1500
# region str 0.0 4
# vehicle_make str 0.0 60
# driver_age float64 0.0 67
# annual_mileage float64 15.9 1265
# prior_claims float64 0.0 5
# policy_tier str 0.0 3
# survey_score float64 26.7 1100
# country str 0.0 1
# claimed int64 0.0 2
Four findings in that one table, before any chart.
policy_id has 1500 unique values across 1500 original rows, which makes it an identifier. Identifiers must be dropped. Leave one in and a tree will happily split on it, memorising the training set.
country has a single value, so it carries no information and should go.
vehicle_make has 60 categories, which is high enough that one-hot encoding adds 60 columns to a 1500-row dataset. Note it now and decide later in this section.
survey_score is 26.7% missing and annual_mileage 15.9%. Both need a decision.
print("duplicate rows:", int(df.duplicated().sum())) # 25
print("constant cols:", [c for c in df.columns if df[c].nunique(dropna=False) <= 1])
# ['country']
print("target rate:", round(float(df["claimed"].mean()), 3)) # 0.401
Duplicates matter more than they look. Twenty-five repeated rows will land on both sides of a train and test split, inflating your score for free.
Missingness is a signal before it is a problem
Never look only at how much is missing. Look at which rows are missing it.
print(df.groupby(df["annual_mileage"].isna())["driver_age"].mean().round(1).to_dict())
# {False: 45.8, True: 75.2}
print(df.groupby(df["annual_mileage"].isna())["claimed"].mean().round(3).to_dict())
# {False: 0.45, True: 0.145}
Policies with missing mileage belong to drivers averaging 75 years old, against 46 when the value is present. Their claim rate is 0.145 against 0.45. The blank is not noise; it carries information about who the customer is.
Two lines of groupby turned a data quality issue into a finding. That comparison is the single highest-value check in this chapter, and it is the one most people skip because the missing-percentage column looks like a sufficient answer.
Relationships with the target
Check how each numeric column moves with the target, both to anticipate which features will matter and to catch anything suspiciously strong.
numeric = df.select_dtypes("number").drop(columns="claimed")
print(numeric.corrwith(df["claimed"]).round(3)
.sort_values(key=abs, ascending=False).to_dict())
# {'driver_age': -0.439, 'prior_claims': 0.406,
# 'annual_mileage': 0.245, 'survey_score': -0.084}
Three columns carry real signal and survey_score carries almost none, which is worth knowing before you spend an afternoon imputing its 400 missing values.
The number to be suspicious of is the one near 0.95. A feature that correlates almost perfectly with the target is usually a consequence of the target rather than a predictor of it, which is the failure described in the main challenges of machine learning. Nothing here is near that, so there is nothing to investigate.
Remember that correlation detects only linear relationships. A feature with a clean U-shaped effect scores near zero, so a low correlation is a reason to look at a scatter plot rather than a reason to drop the column.
The checks, as a list
| Check | What you run | What it tells you |
|---|---|---|
| Shape and dtypes | df.shape, df.dtypes |
Numbers stored as text, dates stored as strings |
| Identifiers | df.nunique() near row count |
Columns to drop before modelling |
| Constants | nunique() <= 1 |
Columns carrying no information |
| Duplicates | df.duplicated().sum() |
Rows that will inflate your score |
| Missingness pattern | groupby(col.isna()) |
Whether blanks are informative |
| Cardinality | nunique() on object columns |
Whether one-hot encoding is affordable |
| Target balance | y.value_counts() |
Whether this is an imbalanced problem |
Run all seven before fitting anything. They take under a minute and they determine most of what the next four chapters will ask you to do.
What this stage deliberately does not do
You now have a repair list: drop policy_id and country, remove 25 duplicates, decide about two columns with missing values, and plan for 60 vehicle makes. You have not fixed any of it, and you should not.
Separating diagnosis from treatment matters because treatment decisions depend on the model you choose, and several of them must happen inside a cross-validation fold rather than on the full dataset. Imputing now, in your exploration notebook, is how test information reaches training. Keep your findings in a list and apply them in a pipeline, as sketched in the project lifecycle.
One exception. If EDA shows the data cannot support the target you defined during target definition, stop and go back rather than proceeding. Column reference details are in the pandas.DataFrame.describe documentation.
Frequently Asked Questions
How is EDA for machine learning different from normal EDA?
Normal EDA answers business questions. Machine learning EDA produces a repair list: which columns are identifiers, which are constant, where the missing values are, how many categories each text column has, and whether the target is balanced. The output is decisions about preparation, not insights.
What should you check first in a new dataset?
Shape, dtypes and a per-column table of missing percentage and unique count. That single table reveals identifier columns, constant columns, high-cardinality text and missingness in one view. Check for duplicate rows next, since they inflate scores by appearing in both training and test splits.
Should you fix missing values during EDA?
No. Record where they are and investigate whether they are informative, then apply the fix inside a pipeline so it refits on each training fold. Imputing in your exploration notebook computes statistics from the whole dataset, which leaks test information into training.
Key Takeaways
- Produce one profile table of dtype, missing percentage and unique count per column, because it exposes identifiers, constants and cardinality problems in a single view.
- Group your target and other features by whether a column is missing, since an informative blank is a finding rather than a defect.
- Count duplicate rows before splitting, as repeated records landing in both halves inflate your score without improving the model.
- Treat a correlation above roughly 0.9 with the target as a warning to investigate, not a strong feature to celebrate.
- Keep your findings as a repair list and apply them inside a pipeline, rather than cleaning the full dataset in your exploration notebook.