Sam Austin on October 9, 2026

How to Detect Bias in Training Data Before You Build a Model

How to Detect Bias in Training Data Before You Build a Model
Contents

Hands sorting stacked cards into uneven piles, the imbalance a representation check looks for

Figure 1: Who's in the pile — and who isn't

You spent two weeks tuning a model, and then a fairness audit showed it rejects one group far more often than another. The model didn't invent that gap. It learned it from your data.

Fixing bias after training costs time, and catching it before training costs an afternoon. This guide shows you five checks you can run on any tabular dataset before you fit a single model. I tested every snippet on a synthetic loan dataset, and I report the real output below. Install what you need with pip install pandas scipy scikit-learn.

Why Data Bias Comes First

A model copies the patterns in its training data. If history treated groups unequally, your labels record that inequality, and your model learns it as signal. Ever wondered why a "neutral" algorithm produces unfair results? Because it never started neutral.

Bias also hides well. Your overall accuracy looks fine, your columns look clean, and nothing crashes. You only find the problem if you look for it on purpose.

Know What You're Hunting

Bias enters training data in a few distinct ways. Name them, and your checks get sharper.

  • Representation bias: some groups appear far less often than they should.
  • Historical bias: the labels reflect past discrimination, like old loan denials.
  • Measurement bias: a feature or label means different things for different groups.
  • Missing-data bias: gaps cluster in certain groups.
  • Proxy bias: an innocent-looking column stands in for a protected attribute.

The five checks below map to these categories. Run them in order.

The Test Dataset

I generated 6,000 synthetic loan applications with problems planted on purpose. Women make up about 30% of rows, historical approvals favor men, women's income values go missing more often, and a zip-code band tracks one racial group. You'll see how each check finds its matching problem.

import numpy as np
import pandas as pd
from scipy.stats import chi2_contingency
from sklearn.ensemble import RandomForestClassifier
from sklearn.model_selection import cross_val_score

One script can run all five checks, so you can change the settings and re-run it against your own data.

Check 1: Representation

Start with the simplest question: who appears in your data, and in what proportion?

print(df["sex"].value_counts(normalize=True).round(3))

My dataset returns about 69% male and 31% female. That gap alone doesn't prove unfairness, because the real population might look the same. But you must compare it against the population your model will serve.

Ask yourself two things:

  1. Does the dataset match the deployment population? A hiring model trained on past hires inherits who got hired before.
  2. Are any groups too small to learn from? Small groups get worse models and noisier metrics.

IMO, this check catches more real-world trouble than any fancy metric, and it takes ten seconds.

Check 2: Label Rates by Group

Now compare how often each group receives the favorable label. This check targets historical bias directly.

rates = df.groupby("sex")["approved"].mean()
print(rates.round(3))
print("impact ratio:", round(rates.min() / rates.max(), 3))

chi2, p, _, _ = chi2_contingency(pd.crosstab(df["sex"], df["approved"]))
print("chi-square p-value:", p)

Here's what my run produced:

  • Female approval rate: 0.312
  • Male approval rate: 0.505
  • Impact ratio: 0.619
  • Chi-square p-value: about 1.9e-43

The impact ratio divides the lower rate by the higher one. US employment guidance uses a 0.8 threshold, known as the four-fifths rule, as a rough screen. FYI, treat it as a flag to investigate, not a legal verdict — the demographic parity vs equalized odds guide explains where the same threshold appears downstream.

The chi-square test tells you whether the gap likely comes from chance. A tiny p-value says the gap is real — it doesn't say why the gap exists, and some gaps have legitimate explanations. A gap that survives scrutiny still needs a defense.

Check 3: Missing Data by Group

Missing values rarely fall at random. When one group's data goes missing more often, imputation and row-dropping can quietly hurt that group.

print(df.drop(columns="sex").isna().groupby(df["sex"]).mean().round(3))

In my dataset, 24.2% of women's income values are missing, versus 0% for men. If you drop incomplete rows, you remove a quarter of the already-small female group. If you fill the gaps with the overall median, you pull women's incomes toward the majority's.

Ask why the data is missing. Sometimes a form field applies differently to different groups, or a data source covers some groups poorly. The pattern of missingness often points to a broken collection process.

Check 4: Proxy Detection

Dropping the sensitive column doesn't remove bias, because other columns can leak it. Zip code, school name, and purchase history often correlate with race, sex, or age.

You can test for leakage directly. Train a quick classifier to predict the protected attribute from the remaining features. If it succeeds, your features carry that information.

X = pd.get_dummies(df.drop(columns=["sex", "approved"]),
                   columns=["race", "zip_band"])
X = X.fillna(X.median())

auc = cross_val_score(
    RandomForestClassifier(n_estimators=100, random_state=0),
    X, (df["sex"] == "female").astype(int),
    cv=5, scoring="roc_auc",
)
print(auc.mean())

An AUC of 0.5 means no leakage, and 1.0 means the features fully reveal the attribute. My run gave 0.687 for sex, a moderate leak through income and credit score. Predicting one racial group from zip band, income, and credit score gave 0.82, a strong proxy.

Keep in mind that I included race in the first test's features, so interpret that number as a rough screen. Rerun the check with only the features you plan to give the model.

Don't panic at a high AUC. Some proxies carry legitimate signal, and you may still need them. The point is to choose to use them, knowing what they reveal.

Check 5: Intersectional Slices

Single-attribute checks miss compounded effects. A group defined by two attributes can face a gap that neither attribute shows alone.

print(df.groupby(["sex", "race"])["approved"]
        .agg(count="size", rate="mean").round(3))

My output shows female applicants in group C with an approval rate of 0.256 across only 219 rows, the lowest of any slice. Male applicants in the same race group sit at 0.499. The smallest, most affected slice rarely announces itself in aggregate numbers.

Watch the counts as closely as the rates. A slice with 30 rows produces a rate you shouldn't trust. Report sample sizes next to every percentage.

What to Do With Your Findings

Finding a problem is only half the job. Match each finding to a response.

Finding Typical response
Under-represented group Collect more data, or reweight carefully
Biased historical labels Question the labels; consider alternative targets
Skewed missingness Fix the collection process before you impute
Strong proxy features Remove, transform, or consciously justify them
Small intersectional slices Report uncertainty; avoid overclaiming

Resist the urge to "fix" everything with a technical trick. Sometimes the honest answer is that the data can't support the decision you want to make. When the numbers say the dataset is usable, the Fairlearn tutorial picks up where this guide stops: measuring the gaps your model actually ships with.

Document What You Found

Write your findings down. A short datasheet for your dataset should record where the data came from, who it covers, what's missing, and which checks you ran. The "Datasheets for Datasets" proposal by Gebru and colleagues popularized this idea, and it pays off when auditors or teammates ask questions six months later.

Include the numbers from each check, even the ugly ones. Future you will thank present you.

Common Mistakes

I've made a few of these myself:

  • Checking only the final model. By then, you've buried the cause.
  • Dropping the sensitive column and calling it done. Proxies survive the drop.
  • Ignoring sample size. A dramatic gap in 20 rows means little.
  • Treating the 0.8 rule as a pass/fail test. It's a screening heuristic.
  • Running checks once. Data changes, so rerun them whenever you refresh it.
CoverBookDescriptionGet it
Datasheets for Datasets the documentation practice this guide's last section builds on. View on Amazon
Cover of “Fairness and Machine Learning” Fairness and Machine Learningby Barocas, Hardt, and Narayanan the taxonomy behind the five categories above. View on Amazon
Cover of “Designing Machine Learning Systems” Designing Machine Learning Systemsby Chip Huyen where data-quality checks belong in the full production loop. View on Amazon

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What are the five checks for bias in training data?

Run them in order: representation (who appears in the data and in what proportion), label rates by group (how often each group receives the favorable label, with an impact ratio and chi-square test), missing data by group (whether gaps cluster in certain groups), proxy detection (train a classifier to predict the protected attribute from your features — high AUC means leakage), and intersectional slices (group by two attributes at once, watching counts as closely as rates). They need only pandas, scipy, and scikit-learn.

What is proxy bias and how do you detect it?

Proxy bias is when an innocent-looking column stands in for a protected attribute: zip code, school name, and purchase history often correlate with race, sex, or age, so dropping the sensitive column doesn't remove the bias. Detect it directly by training a quick classifier to predict the protected attribute from the remaining features and scoring it with cross-validated AUC — 0.5 means no leakage, 1.0 means the features fully reveal the attribute. Then decide deliberately: remove, transform, or consciously justify the proxy.

What does the four-fifths rule mean in a data audit?

The impact ratio divides the lower favorable-label rate by the higher one, and US employment guidance treats values below 0.8 as a flag worth investigating — the same four-fifths screen used with demographic parity metrics. Treat it as a screening heuristic, not a legal verdict: a tiny chi-square p-value says the gap is real, not why it exists, and some gaps have legitimate explanations that still need a defense.

How do you check for intersectional bias?

Group your data by two sensitive attributes at once — for example sex and race — and compare approval counts and rates per slice, because a group defined by two attributes can face a gap that neither attribute shows alone. Watch the counts as closely as the rates: a slice with 30 rows produces a rate you shouldn't trust, so report sample sizes next to every percentage and avoid overclaiming on small cells.

Wrapping This Up

Five checks catch most data-level bias: representation, label rates, missingness, proxies, and intersections. They use tools you already have, and they run in minutes. Bias you find before training costs far less than bias a customer finds after launch.

Open your current dataset this week and run Check 1 and Check 2. If the numbers surprise you, good. A surprise now beats a headline later.

What are You Looking For?

esc