Looking at data to decide what to fix

An analyst explores data to answer questions about the business. You are exploring it to answer one question about the data itself: what will break when I put this into a model? Exploratory data analysis for machine learning is a diagnosis rather than a tour, and it ends with a list of repairs, not a set of charts.

That changes what you look at. Average claim value by region is interesting to a business and irrelevant to you right now. Whether region has four categories or four thousand decides whether one-hot encoding is viable. Whether survey_score is missing at random or missing for a specific kind of customer decides whether imputing it will mislead the model.

Six things to establish, in order. After that the rest of Section 4 does the actual fixing, so resist treating anything yet.

What exploratory data analysis for machine learning has to establish

Start with a profile rather than a plot. One table, every column, three facts each.

import numpy as np
import pandas as pd

rng = np.random.default_rng(12)
n = 1500
df = pd.DataFrame({
    "policy_id": [f"P{100000+i}" for i in range(n)],
    "region": rng.choice(["North", "South", "East", "West"], n, p=[.4, .3, .2, .1]),
    "vehicle_make": rng.choice([f"make_{i}" for i in range(60)], n),
    "driver_age": rng.integers(18, 85, n).astype(float),
    "annual_mileage": rng.gamma(2.0, 6000, n),
    "prior_claims": rng.poisson(0.5, n).astype(float),
    "policy_tier": rng.choice(["bronze", "silver", "gold"], n, p=[.5, .35, .15]),
    "survey_score": rng.normal(7, 1.5, n),
    "country": "IN",
})
df.loc[rng.choice(n, 400, replace=False), "survey_score"] = np.nan
older = df.index[df["driver_age"] > 65]
df.loc[rng.choice(older, int(len(older) * 0.55), replace=False), "annual_mileage"] = np.nan
risk = (0.02 * (80 - df["driver_age"]) + 0.00004 * df["annual_mileage"].fillna(12000)
        + 0.6 * df["prior_claims"])
df["claimed"] = (risk + rng.normal(0, 0.5, n) > 1.55).astype(int)
df = pd.concat([df, df.iloc[:25]], ignore_index=True)   # 25 accidental duplicates

profile = pd.DataFrame({
    "dtype": df.dtypes.astype(str),
    "missing_pct": (df.isna().mean() * 100).round(1),
    "n_unique": df.nunique(),
})
print(profile.to_string())
#                   dtype  missing_pct  n_unique
# policy_id           str          0.0      1500
# region              str          0.0         4
# vehicle_make        str          0.0        60
# driver_age      float64          0.0        67
# annual_mileage  float64         15.9      1265
# prior_claims    float64          0.0         5
# policy_tier         str          0.0         3
# survey_score    float64         26.7      1100
# country             str          0.0         1
# claimed           int64          0.0         2

Four findings in that one table, before any chart.

policy_id has 1500 unique values across 1500 original rows, which makes it an identifier. Identifiers must be dropped. Leave one in and a tree will happily split on it, memorising the training set.

country has a single value, so it carries no information and should go.

vehicle_make has 60 categories, which is high enough that one-hot encoding adds 60 columns to a 1500-row dataset. Note it now and decide later in this section.

survey_score is 26.7% missing and annual_mileage 15.9%. Both need a decision.

print("duplicate rows:", int(df.duplicated().sum()))          # 25
print("constant cols:", [c for c in df.columns if df[c].nunique(dropna=False) <= 1])
# ['country']
print("target rate:", round(float(df["claimed"].mean()), 3))  # 0.401

Duplicates matter more than they look. Twenty-five repeated rows will land on both sides of a train and test split, inflating your score for free.

Missingness is a signal before it is a problem

Never look only at how much is missing. Look at which rows are missing it.

print(df.groupby(df["annual_mileage"].isna())["driver_age"].mean().round(1).to_dict())
# {False: 45.8, True: 75.2}

print(df.groupby(df["annual_mileage"].isna())["claimed"].mean().round(3).to_dict())
# {False: 0.45, True: 0.145}

Policies with missing mileage belong to drivers averaging 75 years old, against 46 when the value is present. Their claim rate is 0.145 against 0.45. The blank is not noise; it carries information about who the customer is.

Two lines of groupby turned a data quality issue into a finding. That comparison is the single highest-value check in this chapter, and it is the one most people skip because the missing-percentage column looks like a sufficient answer.

Relationships with the target

Check how each numeric column moves with the target, both to anticipate which features will matter and to catch anything suspiciously strong.

numeric = df.select_dtypes("number").drop(columns="claimed")
print(numeric.corrwith(df["claimed"]).round(3)
      .sort_values(key=abs, ascending=False).to_dict())
# {'driver_age': -0.439, 'prior_claims': 0.406,
#  'annual_mileage': 0.245, 'survey_score': -0.084}

Three columns carry real signal and survey_score carries almost none, which is worth knowing before you spend an afternoon imputing its 400 missing values.

The number to be suspicious of is the one near 0.95. A feature that correlates almost perfectly with the target is usually a consequence of the target rather than a predictor of it, which is the failure described in the main challenges of machine learning. Nothing here is near that, so there is nothing to investigate.

Remember that correlation detects only linear relationships. A feature with a clean U-shaped effect scores near zero, so a low correlation is a reason to look at a scatter plot rather than a reason to drop the column.

The checks, as a list

Check What you run What it tells you
Shape and dtypes df.shape, df.dtypes Numbers stored as text, dates stored as strings
Identifiers df.nunique() near row count Columns to drop before modelling
Constants nunique() <= 1 Columns carrying no information
Duplicates df.duplicated().sum() Rows that will inflate your score
Missingness pattern groupby(col.isna()) Whether blanks are informative
Cardinality nunique() on object columns Whether one-hot encoding is affordable
Target balance y.value_counts() Whether this is an imbalanced problem

Run all seven before fitting anything. They take under a minute and they determine most of what the next four chapters will ask you to do.

What this stage deliberately does not do

You now have a repair list: drop policy_id and country, remove 25 duplicates, decide about two columns with missing values, and plan for 60 vehicle makes. You have not fixed any of it, and you should not.

Separating diagnosis from treatment matters because treatment decisions depend on the model you choose, and several of them must happen inside a cross-validation fold rather than on the full dataset. Imputing now, in your exploration notebook, is how test information reaches training. Keep your findings in a list and apply them in a pipeline, as sketched in the project lifecycle.

One exception. If EDA shows the data cannot support the target you defined during target definition, stop and go back rather than proceeding. Column reference details are in the pandas.DataFrame.describe documentation.

Frequently Asked Questions

How is EDA for machine learning different from normal EDA?

Normal EDA answers business questions. Machine learning EDA produces a repair list: which columns are identifiers, which are constant, where the missing values are, how many categories each text column has, and whether the target is balanced. The output is decisions about preparation, not insights.

What should you check first in a new dataset?

Shape, dtypes and a per-column table of missing percentage and unique count. That single table reveals identifier columns, constant columns, high-cardinality text and missingness in one view. Check for duplicate rows next, since they inflate scores by appearing in both training and test splits.

Should you fix missing values during EDA?

No. Record where they are and investigate whether they are informative, then apply the fix inside a pipeline so it refits on each training fold. Imputing in your exploration notebook computes statistics from the whole dataset, which leaks test information into training.

Key Takeaways

  • Produce one profile table of dtype, missing percentage and unique count per column, because it exposes identifiers, constants and cardinality problems in a single view.
  • Group your target and other features by whether a column is missing, since an informative blank is a finding rather than a defect.
  • Count duplicate rows before splitting, as repeated records landing in both halves inflate your score without improving the model.
  • Treat a correlation above roughly 0.9 with the target as a warning to investigate, not a strong feature to celebrate.
  • Keep your findings as a repair list and apply them inside a pipeline, rather than cleaning the full dataset in your exploration notebook.