Sam Austin on October 9, 2026

AI Fairness 360 Tutorial: IBM's Bias Detection Toolkit

AI Fairness 360 Tutorial: IBM's Bias Detection Toolkit
Contents

A checklist on a clipboard next to a laptop, the audit AIF360 runs on a model

Figure 1: Seventy metrics, one toolkit, one workflow

Your model passes every accuracy check, and then a reviewer asks, "What's the disparate impact ratio?" You don't have one. AI Fairness 360 (AIF360) gives you that number and about seventy others.

I reach for it when I need metrics and mitigation algorithms in one place. This tutorial walks you through the full loop: load data, measure bias, mitigate it before training, and fix it after training.

What AIF360 Is

AIF360 is IBM's open-source toolkit for detecting and mitigating bias in machine learning. The project describes itself as a library of fairness metrics, explanations for those metrics, and bias mitigation algorithms. It ships in both Python and R, and the maintainers say it's still in development.

Think of it as a fairness workbench rather than a single tool. You get dataset metrics, model metrics, and a long list of mitigation algorithms. Ever wondered how many ways exist to "fix" bias? The README lists a dozen or more, including Reweighing, Disparate Impact Remover, Adversarial Debiasing, and Reject Option Classification.

Install It

The README lists support for Python 3.10 through 3.13. Install the core package like this:

pip install aif360

The metrics work without extras. Some algorithms need optional dependencies, so install them only when you need them:

  • aif360[all] installs everything.
  • aif360[LFR,OptimPreproc] installs selected algorithm dependencies.
  • AdversarialDebiasing needs TensorFlow.
  • OptimPreproc needs CVXPY.

The project recommends a clean conda environment. I agree. Dependency conflicts cause more AIF360 headaches than the library itself.

Learn the Vocabulary

AIF360 uses a few terms you must get right, or every metric flips its sign.

  • Protected attribute: the characteristic you audit, like sex or race.
  • Privileged group: the group that historically receives better outcomes.
  • Unprivileged group: the group that historically receives worse outcomes.
  • Favorable label: the outcome people want, like "loan approved."

You define groups as lists of dictionaries:

privileged = [{"sex": 1}]
unprivileged = [{"sex": 0}]

Every metric then compares the unprivileged group against the privileged one. Swap them accidentally and your sign flips.

Step 1: Build a Dataset

I use a synthetic loan dataset so you don't need to download anything. Men get a small income bump and women get a small credit-score penalty, which creates realistic proxy bias — the same kind of structural gap the synthetic data bias guide covers from the other direction.

import numpy as np
import pandas as pd

rng = np.random.default_rng(42)
n = 8000
sex = rng.binomial(1, 0.55, n)          # 1 = male (privileged), 0 = female
income = rng.normal(55, 15, n) + 6 * sex
credit = rng.normal(650, 60, n) - 20 * (1 - sex)
debt = rng.normal(0.35, 0.1, n)

z = (0.8 * (income - 55) / 15
     + 0.8 * (credit - 650) / 60
     - 0.6 * (debt - 0.35) / 0.1)
label = rng.binomial(1, 1 / (1 + np.exp(-z)))

df = pd.DataFrame({"income": income, "credit_score": credit,
                   "debt_ratio": debt, "sex": sex, "label": label})

Now wrap the DataFrame in AIF360's dataset class and split it three ways:

from aif360.datasets import BinaryLabelDataset

dataset = BinaryLabelDataset(
    df=df,
    label_names=["label"],
    protected_attribute_names=["sex"],
    favorable_label=1,
    unfavorable_label=0,
)

train, valid, test = dataset.split([0.5, 0.75], shuffle=True, seed=0)

That split gives you 50% for training, 25% for validation, and 25% for testing. You need the validation slice later for post-processing.

Step 2: Detect Bias in the Data

Measure the training data before you train anything: biased data produces biased models, and this check tells you early.

from aif360.metrics import BinaryLabelDatasetMetric

metric = BinaryLabelDatasetMetric(
    train, unprivileged_groups=unprivileged, privileged_groups=privileged
)
print("Statistical parity difference:", metric.statistical_parity_difference())
print("Disparate impact:", metric.disparate_impact())

Here's how to read the two numbers:

  • Statistical parity difference: the favorable-outcome rate of the unprivileged group minus the privileged group's rate. Zero means equal rates, and negative values mean the unprivileged group does worse.
  • Disparate impact: the ratio of those two rates. One means equal rates, and values below roughly 0.8 trigger the "four-fifths rule" screening heuristic from US employment law.

FYI, that 0.8 cutoff works as a rough screen, not a legal verdict. Ask counsel before you treat it as a compliance test — the wider demographic parity vs equalized odds guide explains where that threshold comes from.

Step 3: Train a Baseline Model

AIF360 doesn't train models for you. You bring scikit-learn, and AIF360 wraps the predictions.

from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler

scaler = StandardScaler().fit(train.features)
model = LogisticRegression(max_iter=1000)
model.fit(scaler.transform(train.features), train.labels.ravel())

def predict_dataset(model, ds):
    pred = ds.copy(deepcopy=True)
    pred.labels = model.predict(scaler.transform(ds.features)).reshape(-1, 1)
    return pred

test_pred = predict_dataset(model, test)

The helper copies the dataset and swaps in predicted labels. AIF360 compares a "true" dataset against a "predicted" dataset, so you need both.

One catch: train.features includes the protected attribute column here. Drop it from the features if you want a model that never sees sex directly. Proxies like income and credit score will still leak the signal, though.

Step 4: Detect Bias in the Model

Now compare predictions against ground truth with ClassificationMetric:

from aif360.metrics import ClassificationMetric

def audit(true_ds, pred_ds, name):
    cm = ClassificationMetric(
        true_ds, pred_ds,
        unprivileged_groups=unprivileged, privileged_groups=privileged,
    )
    print(f"\n== {name}")
    print("Accuracy:", round(cm.accuracy(), 3))
    print("Disparate impact:", round(cm.disparate_impact(), 3))
    print("Statistical parity diff:", round(cm.statistical_parity_difference(), 3))
    print("Equal opportunity diff:", round(cm.equal_opportunity_difference(), 3))
    print("Average odds diff:", round(cm.average_odds_difference(), 3))

audit(test, test_pred, "baseline")

You now have three families of numbers:

  1. Statistical parity and disparate impact compare selection rates, like demographic parity.
  2. Equal opportunity difference compares true positive rates between groups.
  3. Average odds difference averages the gaps in true and false positive rates.

IMO, reading all of these together beats chasing any single one. They measure different things, and they can disagree.

Step 5: Mitigate Before Training With Reweighing

AIF360 groups mitigation into three stages: pre-processing changes the data, in-processing changes the training, and post-processing changes the predictions. Start with the simplest pre-processing method.

Reweighing assigns a weight to each training example so that the protected attribute and the label become statistically independent. It doesn't change any feature values or labels — it only changes how much each row counts.

from aif360.algorithms.preprocessing import Reweighing

rw = Reweighing(unprivileged_groups=unprivileged, privileged_groups=privileged)
rw.fit(train)
train_rw = rw.transform(train)

The transform call returns a new dataset with updated instance_weights. Pass those weights to your model:

model_rw = LogisticRegression(max_iter=1000)
model_rw.fit(
    scaler.transform(train_rw.features),
    train_rw.labels.ravel(),
    sample_weight=train_rw.instance_weights,
)

audit(test, predict_dataset(model_rw, test), "reweighing")

Notice the pattern: you fit Reweighing on training data only, then apply the model to untouched test data. Never reweigh your test set.

Step 6: Mitigate After Training With Equalized Odds

Post-processing adjusts predictions from a model you've already trained. EqOddsPostprocessing changes some predicted labels to equalize true and false positive rates across groups.

from aif360.algorithms.postprocessing import EqOddsPostprocessing

valid_pred = predict_dataset(model, valid)

eq = EqOddsPostprocessing(
    unprivileged_groups=unprivileged, privileged_groups=privileged, seed=0
)
eq.fit(valid, valid_pred)

test_pred_eq = eq.predict(test_pred)
audit(test, test_pred_eq, "equalized odds postprocessing")

You fit on the validation set and apply to the test set. Fitting on the same data you evaluate on inflates your results.

Post-processing carries a legal and ethical catch: it treats groups differently at decision time, and some jurisdictions restrict that. Check the rules before you ship it.

Compare the Approaches

Run all three audits and line them up. Ask three questions:

  1. Did the fairness gaps shrink? Compare each metric against the baseline.
  2. What did accuracy cost? Fairness usually trades off against accuracy.
  3. Do the metrics agree? Fixing one gap sometimes widens another.
Reweighing EqOddsPostprocessing
Stage Pre-processing Post-processing
Changes Training sample weights Predicted labels
Retrains model? Yes No
Needs protected attribute at prediction? No Yes

If you can retrain and can't use protected attributes at decision time, Reweighing is the safer default. If you inherit a finished model, post-processing gives the quickest path.

Common Mistakes

I've made most of these myself:

  • Swapping privileged and unprivileged groups. Double-check your dictionaries, because the sign of every metric depends on them.
  • Evaluating on training data. Always audit on held-out data.
  • Reweighing the test set. Mitigate training data only.
  • Fixing the metric after seeing results. Choose your fairness definition first.
  • Ignoring intersections. A model can look fair on sex and on race separately and still fail for specific combinations.
CoverBookDescriptionGet it
Cover of “Fairness and Machine Learning” Fairness and Machine Learningby Barocas, Hardt, and Narayanan the theory behind every metric in this tutorial. View on Amazon
Cover of “Designing Machine Learning Systems” Designing Machine Learning Systemsby Chip Huyen places fairness checks inside the full production loop this tutorial plugs into. View on Amazon
Cover of “Interpretable Machine Learning” Interpretable Machine Learningby Christoph Molnar pairs measurement with the explanation methods from the rest of this series. View on Amazon

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What is AI Fairness 360 (AIF360)?

AIF360 is IBM's open-source toolkit for detecting and mitigating bias in machine learning: a library of fairness metrics, explanations for those metrics, and bias mitigation algorithms, available in both Python and R. Think of it as a fairness workbench rather than a single tool — you get dataset metrics, model metrics, and a dozen or more mitigation algorithms such as Reweighing, Disparate Impact Remover, Adversarial Debiasing, and Reject Option Classification.

What is the difference between statistical parity difference and disparate impact?

Statistical parity difference is subtraction: the favorable-outcome rate of the unprivileged group minus the privileged group's rate, where zero means equal rates and negative values mean the unprivileged group does worse. Disparate impact is division: the ratio of those two rates, where one means equal rates and values below roughly 0.8 trigger the four-fifths rule screening heuristic from US employment law. The 0.8 cutoff is a rough screen, not a legal verdict — ask counsel before treating it as a compliance test.

What are privileged and unprivileged groups in AIF360?

The protected attribute is the characteristic you audit, like sex or race; the privileged group is the one that historically receives better outcomes, and the unprivileged group the one that historically receives worse. You define both as lists of dictionaries, such as privileged = [{"sex": 1}] and unprivileged = [{"sex": 0}], and every metric then compares the unprivileged group against the privileged one. Swap them accidentally and the sign of every metric flips.

How do you mitigate bias with AIF360?

AIF360 groups mitigation into three stages: pre-processing changes the data (Reweighing assigns instance weights so the protected attribute and label become statistically independent), in-processing changes the training, and post-processing changes the predictions (EqOddsPostprocessing equalizes true and false positive rates across groups). Fit any mitigation on training data only, fit post-processing on a validation slice, and always audit on held-out test data — never reweigh the test set, and never fit the mitigation on the data you evaluate on.

Wrapping This Up

AIF360 gives you a full workflow: measure the data, measure the model, mitigate, and measure again. Its breadth is its biggest strength, and its steep vocabulary is its biggest cost. Spend an hour on the group definitions and the rest gets easy.

Run this tutorial on your own dataset this week, and read the dataset metrics before you train anything. The data often tells you the ending before the model even starts.

What are You Looking For?

esc