Sam Austin on October 9, 2026

Fairlearn Tutorial: Detect and Mitigate Bias in Python

Fairlearn Tutorial: Detect and Mitigate Bias in Python
Contents

A magnifying glass over a printed spreadsheet, inspecting the numbers a fairness audit looks at

Figure 1: The per-group table tells a more honest story than the headline accuracy

Your loan model scores well on accuracy, then someone checks approval rates by gender. The gap is big, and now you need to explain it, measure it, and fix it. Fairlearn does all three in a few dozen lines of Python.

I use it whenever a model touches decisions about people. This tutorial walks you through the full loop: build a biased baseline, detect the bias, and mitigate it two different ways.

What Fairlearn Does

Fairlearn is an open-source Python toolkit for assessing and improving fairness in machine learning. It has two halves. The metrics half shows you where your model treats groups differently. The mitigation half reduces those gaps.

It plugs into scikit-learn, so you don't rewrite your pipeline. If you can call .fit() and .predict(), you can use it.

Setup

Install the three packages you need:

pip install fairlearn scikit-learn pandas

Then import everything this tutorial uses:

import numpy as np
import pandas as pd
from sklearn.linear_model import LogisticRegression
from sklearn.model_selection import train_test_split
from sklearn.metrics import accuracy_score
from fairlearn.metrics import (
    MetricFrame, selection_rate, true_positive_rate, false_positive_rate,
    demographic_parity_difference, equalized_odds_difference,
)
from fairlearn.reductions import ExponentiatedGradient, EqualizedOdds
from fairlearn.postprocessing import ThresholdOptimizer

Step 1: Build a Dataset With Built-In Bias

Real data carries messy history, so I simulate that. This synthetic loan dataset gives men a small income bump and women a small credit-score penalty. Those shifts mimic the structural gaps you see in real lending data.

rng = np.random.default_rng(42)
n = 8000
gender = rng.choice(["female", "male"], n, p=[0.45, 0.55])
income = rng.normal(55, 15, n) + np.where(gender == "male", 6, 0)
credit = rng.normal(650, 60, n) - np.where(gender == "female", 20, 0)
debt = rng.normal(0.35, 0.1, n)

z = (0.8 * (income - 55) / 15
     + 0.8 * (credit - 650) / 60
     - 0.6 * (debt - 0.35) / 0.1)
repaid = rng.binomial(1, 1 / (1 + np.exp(-z)))

X = pd.DataFrame({"income": income, "credit_score": credit, "debt_ratio": debt})
y = pd.Series(repaid)
A = pd.Series(gender, name="gender")

Notice that the model never sees gender as a feature — and ever wondered whether dropping the sensitive column makes a model fair? It doesn't, because income and credit score act as proxies for gender here.

Step 2: Train a Baseline Model

Split the data first, and keep the sensitive attribute alongside it:

X_tr, X_te, y_tr, y_te, A_tr, A_te = train_test_split(
    X, y, A, test_size=0.3, random_state=0, stratify=A
)

model = LogisticRegression(max_iter=1000)
model.fit(X_tr, y_tr)
pred = model.predict(X_te)

Always evaluate fairness on the held-out test set. Training-set numbers flatter your model.

Step 3: Detect Bias With MetricFrame

MetricFrame is the center of Fairlearn's metrics: it computes any metric overall and broken down by group. I wrap it in a small helper so I can reuse it after mitigation:

def report(name, y_pred):
    mf = MetricFrame(
        metrics={
            "accuracy": accuracy_score,
            "selection_rate": selection_rate,
            "tpr": true_positive_rate,
            "fpr": false_positive_rate,
        },
        y_true=y_te, y_pred=y_pred, sensitive_features=A_te,
    )
    print("\n==", name)
    print(mf.by_group.round(3))
    print("overall accuracy:", round(accuracy_score(y_te, y_pred), 3))
    print("DP difference:", round(
        demographic_parity_difference(y_te, y_pred, sensitive_features=A_te), 3))
    print("EO difference:", round(
        equalized_odds_difference(y_te, y_pred, sensitive_features=A_te), 3))

report("baseline", pred)

Here's what each column tells you:

  • selection_rate: the share of each group that gets a positive prediction (approval).
  • tpr: among people who would repay, how many the model approves.
  • fpr: among people who would default, how many the model approves.
  • accuracy: the classic score, now split by group.

Read the two summary numbers carefully. Demographic parity difference measures the largest gap in selection rates between groups, and zero means every group gets approved at the same rate. Equalized odds difference measures the larger of the gaps in true positive rate and false positive rate, and zero means the model makes errors at equal rates across groups.

Run the helper, then look at the gaps. In a dataset built like this one, I expect the group with the income and credit advantage to show the higher selection rate. Check your own output instead of trusting my prediction.

Step 4: Mitigate With ExponentiatedGradient

Fairlearn offers two main mitigation families. The first, in-processing, trains the model under a fairness constraint. ExponentiatedGradient implements the reductions approach: it wraps your estimator, retrains it several times with adjusted sample weights, and combines the results.

mitigator = ExponentiatedGradient(
    LogisticRegression(max_iter=1000),
    constraints=EqualizedOdds(),
    eps=0.01,
)
mitigator.fit(X_tr, y_tr, sensitive_features=A_tr)
report("ExponentiatedGradient + EqualizedOdds", mitigator.predict(X_te))

Three details matter here:

  1. Pick the constraint to match your goal. EqualizedOdds() targets error-rate gaps. Swap in DemographicParity() if you care about selection rates instead.
  2. eps sets the tolerance. Smaller values demand tighter fairness and usually cost more accuracy.
  3. You need sensitive_features at training time, but not necessarily at prediction time.

The result is a randomized ensemble, so predictions can vary slightly between calls. Set a seed in your estimator if you need strict reproducibility.

Step 5: Mitigate With ThresholdOptimizer

The second family, post-processing, leaves your trained model alone. ThresholdOptimizer learns group-specific decision thresholds on top of the model's scores.

optimizer = ThresholdOptimizer(
    estimator=model,
    constraints="equalized_odds",
    predict_method="predict_proba",
    prefit=True,
)
optimizer.fit(X_tr, y_tr, sensitive_features=A_tr)

fair_pred = optimizer.predict(X_te, sensitive_features=A_te, random_state=0)
report("ThresholdOptimizer equalized_odds", fair_pred)

This approach has a catch: it needs the sensitive attribute at prediction time. Many organizations can't use protected attributes at decision time, and some jurisdictions restrict group-specific thresholds. Talk to legal before you ship this.

FYI, newer Fairlearn releases have been changing how prefit works, so check the current docs if your version warns you about it.

Step 6: Compare the Results

Run all three reports and compare them side by side. You want to answer three questions:

  1. Did the fairness gaps shrink? Compare the DP and EO differences against the baseline.
  2. What did accuracy cost? Compare overall accuracy across all three models.
  3. Did the group-level rates move the way you expected? Look at tpr and fpr per group.

IMO, this comparison matters more than any single number: a fairness gain that costs three accuracy points might be a bargain or a dealbreaker, depending on your use case. You make that call, not the library.

Choosing Between the Two Approaches

ExponentiatedGradient ThresholdOptimizer
Type In-processing Post-processing
Retrains the model? Yes No
Needs sensitive attribute at prediction? No Yes
Best for New models you control Existing models you can't retrain

If you can retrain and can't use protected attributes at decision time, pick ExponentiatedGradient. If you inherit a finished model, ThresholdOptimizer gives the quickest path.

Common Pitfalls

I've tripped over most of these:

  • Measuring on training data. Always use held-out data.
  • Choosing the metric after seeing results. Decide your fairness definition first — that pattern is dissected in the demographic parity vs equalized odds guide.
  • Ignoring small groups. Tiny samples produce noisy metrics, so report uncertainty.
  • Skipping intersections. A model can look fair on gender and race separately and still fail for specific combinations. MetricFrame accepts multiple sensitive columns, so check them.
  • Treating the numbers as certification. Metrics describe patterns. They don't prove a model is "fair." Metrics also decay — keep watching them the way you'd watch drift through model monitoring in production.
CoverBookDescriptionGet it
Cover of “Fairness and Machine Learning” Fairness and Machine Learningby Barocas, Hardt, and Narayanan the theory behind the constraints you pass to Fairlearn. View on Amazon
Cover of “Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow” Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlowby Aurélien Géron the pipeline fundamentals this tutorial assumes. View on Amazon
Cover of “Interpretable Machine Learning” Interpretable Machine Learningby Christoph Molnar measurement and explanation methods that pair with fairness metrics. View on Amazon

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What is Fairlearn and what is it used for?

Fairlearn is an open-source Python toolkit for assessing and improving fairness in machine learning, with two halves: metrics that show where your model treats groups differently, and mitigation algorithms that reduce those gaps. It plugs into scikit-learn, so if you can call .fit() and .predict(), you can use it without rewriting your pipeline. The typical loop is measure with MetricFrame, mitigate with ExponentiatedGradient or ThresholdOptimizer, then measure again.

Does dropping the sensitive feature make a model fair?

No. Even when gender is not a feature, income and credit score can act as proxies for it, so the model reproduces the same gaps. Fairness has to be measured on predictions with the sensitive attribute attached — split the data keeping the sensitive attribute alongside it, and evaluate on the held-out test set, since training-set numbers flatter your model.

What is the difference between ExponentiatedGradient and ThresholdOptimizer?

ExponentiatedGradient is in-processing: it retrains the model under a fairness constraint by wrapping your estimator, retraining with adjusted sample weights, and combining the results, so it needs the sensitive attribute at training time but not prediction time. ThresholdOptimizer is post-processing: it leaves the trained model alone and learns group-specific decision thresholds on top of its scores, but needs the sensitive attribute at prediction time. Retrain when you can and can't use protected attributes at decision time; otherwise post-processing is the quickest path.

How do you measure fairness gaps with Fairlearn?

Use MetricFrame to compute any metric overall and broken down by group — accuracy, selection rate, true positive rate, and false positive rate — then summarize with demographic_parity_difference and equalized_odds_difference, where zero means perfect fairness. Run these on held-out data, report uncertainty for small groups, and check intersections with multiple sensitive columns, because a model can look fair on gender and race separately while failing for specific combinations.

Wrapping This Up

Fairlearn gives you a clear loop: measure with MetricFrame, mitigate with ExponentiatedGradient or ThresholdOptimizer, then measure again. The code stays short, but the decisions behind it don't — you still choose the fairness definition, the accuracy tradeoff, and the legal approach.

Here's your next step. Copy the code, run it on your own model, and look at the per-group table before you look at any summary number. The gaps usually tell a more honest story than the headline accuracy.

What are You Looking For?

esc