Contents
Figure 1: Seventy metrics, one toolkit, one workflow
Your model passes every accuracy check, and then a reviewer asks, "What's the disparate impact ratio?" You don't have one. AI Fairness 360 (AIF360) gives you that number and about seventy others.
I reach for it when I need metrics and mitigation algorithms in one place. This tutorial walks you through the full loop: load data, measure bias, mitigate it before training, and fix it after training.
What AIF360 Is
AIF360 is IBM's open-source toolkit for detecting and mitigating bias in machine learning. The project describes itself as a library of fairness metrics, explanations for those metrics, and bias mitigation algorithms. It ships in both Python and R, and the maintainers say it's still in development.
Think of it as a fairness workbench rather than a single tool. You get dataset metrics, model metrics, and a long list of mitigation algorithms. Ever wondered how many ways exist to "fix" bias? The README lists a dozen or more, including Reweighing, Disparate Impact Remover, Adversarial Debiasing, and Reject Option Classification.
Install It
The README lists support for Python 3.10 through 3.13. Install the core package like this:
pip install aif360
The metrics work without extras. Some algorithms need optional dependencies, so install them only when you need them:
aif360[all]installs everything.aif360[LFR,OptimPreproc]installs selected algorithm dependencies.- AdversarialDebiasing needs TensorFlow.
- OptimPreproc needs CVXPY.
The project recommends a clean conda environment. I agree. Dependency conflicts cause more AIF360 headaches than the library itself.
Learn the Vocabulary
AIF360 uses a few terms you must get right, or every metric flips its sign.
- Protected attribute: the characteristic you audit, like sex or race.
- Privileged group: the group that historically receives better outcomes.
- Unprivileged group: the group that historically receives worse outcomes.
- Favorable label: the outcome people want, like "loan approved."
You define groups as lists of dictionaries:
privileged = [{"sex": 1}]
unprivileged = [{"sex": 0}]
Every metric then compares the unprivileged group against the privileged one. Swap them accidentally and your sign flips.
Step 1: Build a Dataset
I use a synthetic loan dataset so you don't need to download anything. Men get a small income bump and women get a small credit-score penalty, which creates realistic proxy bias — the same kind of structural gap the synthetic data bias guide covers from the other direction.
import numpy as np
import pandas as pd
rng = np.random.default_rng(42)
n = 8000
sex = rng.binomial(1, 0.55, n) # 1 = male (privileged), 0 = female
income = rng.normal(55, 15, n) + 6 * sex
credit = rng.normal(650, 60, n) - 20 * (1 - sex)
debt = rng.normal(0.35, 0.1, n)
z = (0.8 * (income - 55) / 15
+ 0.8 * (credit - 650) / 60
- 0.6 * (debt - 0.35) / 0.1)
label = rng.binomial(1, 1 / (1 + np.exp(-z)))
df = pd.DataFrame({"income": income, "credit_score": credit,
"debt_ratio": debt, "sex": sex, "label": label})
Now wrap the DataFrame in AIF360's dataset class and split it three ways:
from aif360.datasets import BinaryLabelDataset
dataset = BinaryLabelDataset(
df=df,
label_names=["label"],
protected_attribute_names=["sex"],
favorable_label=1,
unfavorable_label=0,
)
train, valid, test = dataset.split([0.5, 0.75], shuffle=True, seed=0)
That split gives you 50% for training, 25% for validation, and 25% for testing. You need the validation slice later for post-processing.
Step 2: Detect Bias in the Data
Measure the training data before you train anything: biased data produces biased models, and this check tells you early.
from aif360.metrics import BinaryLabelDatasetMetric
metric = BinaryLabelDatasetMetric(
train, unprivileged_groups=unprivileged, privileged_groups=privileged
)
print("Statistical parity difference:", metric.statistical_parity_difference())
print("Disparate impact:", metric.disparate_impact())
Here's how to read the two numbers:
- Statistical parity difference: the favorable-outcome rate of the unprivileged group minus the privileged group's rate. Zero means equal rates, and negative values mean the unprivileged group does worse.
- Disparate impact: the ratio of those two rates. One means equal rates, and values below roughly 0.8 trigger the "four-fifths rule" screening heuristic from US employment law.
FYI, that 0.8 cutoff works as a rough screen, not a legal verdict. Ask counsel before you treat it as a compliance test — the wider demographic parity vs equalized odds guide explains where that threshold comes from.
Step 3: Train a Baseline Model
AIF360 doesn't train models for you. You bring scikit-learn, and AIF360 wraps the predictions.
from sklearn.linear_model import LogisticRegression
from sklearn.preprocessing import StandardScaler
scaler = StandardScaler().fit(train.features)
model = LogisticRegression(max_iter=1000)
model.fit(scaler.transform(train.features), train.labels.ravel())
def predict_dataset(model, ds):
pred = ds.copy(deepcopy=True)
pred.labels = model.predict(scaler.transform(ds.features)).reshape(-1, 1)
return pred
test_pred = predict_dataset(model, test)
The helper copies the dataset and swaps in predicted labels. AIF360 compares a "true" dataset against a "predicted" dataset, so you need both.
One catch: train.features includes the protected attribute column here. Drop it from the features if you want a model that never sees sex directly. Proxies like income and credit score will still leak the signal, though.
Step 4: Detect Bias in the Model
Now compare predictions against ground truth with ClassificationMetric:
from aif360.metrics import ClassificationMetric
def audit(true_ds, pred_ds, name):
cm = ClassificationMetric(
true_ds, pred_ds,
unprivileged_groups=unprivileged, privileged_groups=privileged,
)
print(f"\n== {name}")
print("Accuracy:", round(cm.accuracy(), 3))
print("Disparate impact:", round(cm.disparate_impact(), 3))
print("Statistical parity diff:", round(cm.statistical_parity_difference(), 3))
print("Equal opportunity diff:", round(cm.equal_opportunity_difference(), 3))
print("Average odds diff:", round(cm.average_odds_difference(), 3))
audit(test, test_pred, "baseline")
You now have three families of numbers:
- Statistical parity and disparate impact compare selection rates, like demographic parity.
- Equal opportunity difference compares true positive rates between groups.
- Average odds difference averages the gaps in true and false positive rates.
IMO, reading all of these together beats chasing any single one. They measure different things, and they can disagree.
Step 5: Mitigate Before Training With Reweighing
AIF360 groups mitigation into three stages: pre-processing changes the data, in-processing changes the training, and post-processing changes the predictions. Start with the simplest pre-processing method.
Reweighing assigns a weight to each training example so that the protected attribute and the label become statistically independent. It doesn't change any feature values or labels — it only changes how much each row counts.
from aif360.algorithms.preprocessing import Reweighing
rw = Reweighing(unprivileged_groups=unprivileged, privileged_groups=privileged)
rw.fit(train)
train_rw = rw.transform(train)
The transform call returns a new dataset with updated instance_weights. Pass those weights to your model:
model_rw = LogisticRegression(max_iter=1000)
model_rw.fit(
scaler.transform(train_rw.features),
train_rw.labels.ravel(),
sample_weight=train_rw.instance_weights,
)
audit(test, predict_dataset(model_rw, test), "reweighing")
Notice the pattern: you fit Reweighing on training data only, then apply the model to untouched test data. Never reweigh your test set.
Step 6: Mitigate After Training With Equalized Odds
Post-processing adjusts predictions from a model you've already trained. EqOddsPostprocessing changes some predicted labels to equalize true and false positive rates across groups.
from aif360.algorithms.postprocessing import EqOddsPostprocessing
valid_pred = predict_dataset(model, valid)
eq = EqOddsPostprocessing(
unprivileged_groups=unprivileged, privileged_groups=privileged, seed=0
)
eq.fit(valid, valid_pred)
test_pred_eq = eq.predict(test_pred)
audit(test, test_pred_eq, "equalized odds postprocessing")
You fit on the validation set and apply to the test set. Fitting on the same data you evaluate on inflates your results.
Post-processing carries a legal and ethical catch: it treats groups differently at decision time, and some jurisdictions restrict that. Check the rules before you ship it.
Compare the Approaches
Run all three audits and line them up. Ask three questions:
- Did the fairness gaps shrink? Compare each metric against the baseline.
- What did accuracy cost? Fairness usually trades off against accuracy.
- Do the metrics agree? Fixing one gap sometimes widens another.
| Reweighing | EqOddsPostprocessing | |
|---|---|---|
| Stage | Pre-processing | Post-processing |
| Changes | Training sample weights | Predicted labels |
| Retrains model? | Yes | No |
| Needs protected attribute at prediction? | No | Yes |
If you can retrain and can't use protected attributes at decision time, Reweighing is the safer default. If you inherit a finished model, post-processing gives the quickest path.
Common Mistakes
I've made most of these myself:
- Swapping privileged and unprivileged groups. Double-check your dictionaries, because the sign of every metric depends on them.
- Evaluating on training data. Always audit on held-out data.
- Reweighing the test set. Mitigate training data only.
- Fixing the metric after seeing results. Choose your fairness definition first.
- Ignoring intersections. A model can look fair on sex and on race separately and still fail for specific combinations.
Recommended Books
| Cover | Book | Description | Get it |
|---|---|---|---|
![]() |
Fairness and Machine Learning | the theory behind every metric in this tutorial. | View on Amazon |
![]() |
Designing Machine Learning Systems | places fairness checks inside the full production loop this tutorial plugs into. | View on Amazon |
![]() |
Interpretable Machine Learning | pairs measurement with the explanation methods from the rest of this series. | View on Amazon |
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What is AI Fairness 360 (AIF360)?
AIF360 is IBM's open-source toolkit for detecting and mitigating bias in machine learning: a library of fairness metrics, explanations for those metrics, and bias mitigation algorithms, available in both Python and R. Think of it as a fairness workbench rather than a single tool — you get dataset metrics, model metrics, and a dozen or more mitigation algorithms such as Reweighing, Disparate Impact Remover, Adversarial Debiasing, and Reject Option Classification.
What is the difference between statistical parity difference and disparate impact?
Statistical parity difference is subtraction: the favorable-outcome rate of the unprivileged group minus the privileged group's rate, where zero means equal rates and negative values mean the unprivileged group does worse. Disparate impact is division: the ratio of those two rates, where one means equal rates and values below roughly 0.8 trigger the four-fifths rule screening heuristic from US employment law. The 0.8 cutoff is a rough screen, not a legal verdict — ask counsel before treating it as a compliance test.
What are privileged and unprivileged groups in AIF360?
The protected attribute is the characteristic you audit, like sex or race; the privileged group is the one that historically receives better outcomes, and the unprivileged group the one that historically receives worse. You define both as lists of dictionaries, such as privileged = [{"sex": 1}] and unprivileged = [{"sex": 0}], and every metric then compares the unprivileged group against the privileged one. Swap them accidentally and the sign of every metric flips.
How do you mitigate bias with AIF360?
AIF360 groups mitigation into three stages: pre-processing changes the data (Reweighing assigns instance weights so the protected attribute and label become statistically independent), in-processing changes the training, and post-processing changes the predictions (EqOddsPostprocessing equalizes true and false positive rates across groups). Fit any mitigation on training data only, fit post-processing on a validation slice, and always audit on held-out test data — never reweigh the test set, and never fit the mitigation on the data you evaluate on.
Wrapping This Up
AIF360 gives you a full workflow: measure the data, measure the model, mitigate, and measure again. Its breadth is its biggest strength, and its steep vocabulary is its biggest cost. Spend an hour on the group definitions and the rest gets easy.
Run this tutorial on your own dataset this week, and read the dataset metrics before you train anything. The data often tells you the ending before the model even starts.


