Contents
Figure 1: 94% accuracy looks great until someone slices it by group
Your model hits 94% accuracy, and your stakeholders cheer. Then someone slices the results by gender, and the cheering stops. Approval rates differ sharply, and nobody can explain why.
I've watched this scene play out, and it always ends with the same question: "Which fairness metric should we use?" This guide compares the two most popular answers, demographic parity and equalized odds, so you can pick one on purpose instead of by accident.
Why Fairness Metrics Exist
Accuracy hides problems. A model can score well overall and still fail one group badly, because the majority group dominates the average. Fairness metrics force you to look at each group separately.
Think of them as smoke detectors. They don't fix the fire, but they tell you where to look. And just like smoke detectors, different types catch different problems.
The Setup: Four Terms You Need
Both metrics build on a few basic ideas. Get these straight and everything else gets easier.
- Sensitive attribute (A): the characteristic you protect, like gender, race, or age.
- Prediction (Ŷ): what your model outputs, such as "approve" or "deny."
- True outcome (Y): what actually happened, such as "repaid the loan" or "defaulted."
- Group: everyone who shares the same value of the sensitive attribute.
Now you can compare the metrics without getting lost.
Demographic Parity Explained
Demographic parity asks one simple question: does each group receive positive predictions at the same rate? If your model approves 40% of group A, it should approve about 40% of group B.
Notice what this metric ignores: it never looks at the true outcome. It only compares prediction rates across groups.
How you measure it. You compute the positive prediction rate for each group, then compare them. Two common summaries work well:
- Difference: subtract one group's rate from another's. Zero means perfect parity.
- Ratio: divide the lower rate by the higher rate. The US "four-fifths rule" in employment settings uses a ratio threshold of 0.8, which many teams borrow as a rough guide.
FYI, that 0.8 threshold works as a screening heuristic, not a legal verdict. Ask a lawyer before you treat it as a compliance test.
Where Demographic Parity Shines
I like demographic parity when the goal is equal access. Think of ad delivery, outreach, or shortlisting candidates for interviews. In those cases, you care about who gets a chance, not who ultimately succeeds.
It also needs no ground-truth labels at test time, which helps when labels come late or carry bias themselves.
Where It Falls Short
Here's the catch: if the real base rates differ between groups, forcing equal prediction rates can force bad predictions. Imagine two groups where one genuinely repays loans at a higher rate. To hit parity, your model must approve unqualified applicants in one group or reject qualified applicants in the other.
Ever wondered why some people hate this metric? That tradeoff explains it: a model can achieve perfect parity by acting randomly, which helps nobody.
Equalized Odds Explained
Equalized odds asks a smarter question: does the model make errors at the same rates for every group? Specifically, it requires two things at once.
- Equal true positive rates: among people who truly qualify, each group gets approved at the same rate.
- Equal false positive rates: among people who truly don't qualify, each group gets wrongly approved at the same rate.
Hardt, Price, and Srebro introduced this formulation in 2016, and it quickly became a standard. It conditions on the true outcome, so it respects real differences between groups.
The Lighter Cousin: Equal Opportunity
Sometimes you only care about one error type. Equal opportunity requires equal true positive rates and drops the false positive requirement. Use it when missing a qualified person hurts more than approving an unqualified one.
Hiring and scholarship decisions often fit this pattern. You want qualified people to succeed regardless of group.
Where Equalized Odds Shines
I prefer equalized odds in high-stakes decisions like lending, medical triage, and risk assessment. It asks whether the model treats equally qualified people equally. That framing matches how most people think about fairness.
It also avoids the random-guessing loophole. A useless model can't satisfy equalized odds without matching error rates, which requires some real signal.
Where It Falls Short
This metric needs reliable ground-truth labels, and labels often carry bias. If past decisions discriminated, the "true" outcomes in your data inherit that discrimination. Equalized odds then locks in the old bias with a fairness stamp on top.
It also demands more of your model. Matching two error rates across groups usually costs some accuracy.
Head-to-Head Comparison
| Feature | Demographic Parity | Equalized Odds |
|---|---|---|
| Core question | Same positive rate per group? | Same error rates per group? |
| Uses true labels? | No | Yes |
| Handles different base rates? | Poorly | Well |
| Best for | Equal access, outreach | Lending, healthcare, risk scoring |
| Biggest risk | Forces bad predictions | Inherits biased labels |
IMO, this table captures 80% of the decision. The remaining 20% depends on context.
You Can't Have Everything
Here's the uncomfortable truth: researchers proved that several fairness definitions cannot hold simultaneously when base rates differ between groups. Kleinberg and colleagues showed this in 2016, and Chouldechova reached a similar conclusion in 2017.
The famous example comes from the COMPAS recidivism debate. ProPublica argued the tool treated groups unfairly because false positive rates differed. The vendor argued the tool stayed fair because its risk scores meant the same thing across groups. Both sides used valid metrics and reached opposite conclusions.
So who wins that debate? Neither, because the metrics disagree by design. You choose which fairness you value most, then defend that choice openly.
How to Choose Between Them
Skip the theory spiral. Walk through these questions instead.
- What decision does the model support? Access decisions favor demographic parity. Merit-based decisions favor equalized odds.
- How trustworthy are your labels? Biased labels weaken equalized odds.
- Which error hurts more? If false negatives hurt most, try equal opportunity.
- What do regulators and stakeholders expect? Some domains publish clear guidance.
- Can you explain your choice in plain language? If you can't, rethink it.
Document the answers. Future you, and your auditors, will thank you.
Measuring and Fixing It in Practice
You don't need to build metrics from scratch. Open-source libraries do the heavy lifting — the interpretability tools roundup covers adjacent ground, and fairness tooling is more mature than most teams expect.
- Fairlearn offers functions like
demographic_parity_differenceandequalized_odds_difference, plus mitigation tools. - AIF360 from IBM provides a broad toolkit of metrics and bias-mitigation algorithms.
A quick Fairlearn check looks like this:
from fairlearn.metrics import (
demographic_parity_difference,
equalized_odds_difference,
)
dp = demographic_parity_difference(y_true, y_pred, sensitive_features=group)
eo = equalized_odds_difference(y_true, y_pred, sensitive_features=group)
print(dp, eo)
Both functions return a gap, and zero means perfect fairness. Run them on a held-out test set, never on your training data.
Three Ways to Reduce the Gap
Mitigation techniques fall into three families.
- Pre-processing: reweight or resample the training data before you train.
- In-processing: add a fairness constraint while the model trains.
- Post-processing: adjust decision thresholds per group after training, which directly targets equalized odds.
Post-processing often gives the fastest win, because you don't retrain anything. Check the legal rules first, though, since group-specific thresholds raise questions in some jurisdictions. And if you're handing individuals concrete changes after the decision — the kind of guidance the counterfactual explanations guide covers — remember that fairness costs differ across groups: one group may need far larger changes than another, which is disparity worth auditing in its own right.
Common Mistakes
Picking a metric after seeing results: that's fishing, not fairness. The rest come from painful experience.
- Ignoring intersections. A model can look fair for gender and for race separately, yet fail for specific combinations.
- Treating one number as proof. Fairness metrics describe patterns. They don't certify anything.
- Skipping small groups. Tiny samples produce noisy metrics, so report confidence intervals.
- Forgetting the feedback loop. Today's predictions shape tomorrow's training data.
Recommended Books
| Cover | Book | Description | Get it |
|---|---|---|---|
![]() |
Fairness and Machine Learning | the textbook behind the impossibility results in this guide, free online and worth the buy for the worked examples. | View on Amazon |
![]() |
Weapons of Math Destruction | the COMPAS-style failure modes in narrative form. | View on Amazon |
![]() |
Interpretable Machine Learning | connects fairness measurement with the explanation methods from the rest of this series. | View on Amazon |
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What is the difference between demographic parity and equalized odds?
Demographic parity compares outcomes: each group should receive positive predictions at the same rate, and it ignores true labels entirely. Equalized odds compares errors: true positive rates and false positive rates must both match across groups, which requires ground-truth labels. Parity fits equal-access decisions like outreach and interview shortlisting; equalized odds fits merit-based, high-stakes decisions like lending, triage, and risk scoring.
When should you use demographic parity instead of equalized odds?
Use demographic parity when the goal is equal access rather than equal accuracy — ad delivery, outreach, or shortlisting candidates — especially when labels arrive late or carry bias themselves. Be aware of the tradeoff: if real base rates differ between groups, forcing equal prediction rates means approving unqualified applicants in one group or rejecting qualified applicants in the other, and a random-guessing model can achieve perfect parity.
Can a model satisfy both fairness metrics at the same time?
Not when base rates differ between groups. Kleinberg and colleagues showed in 2016, and Chouldechova confirmed in 2017, that several fairness definitions cannot hold simultaneously — the same result behind the COMPAS debate, where ProPublica and the vendor each used a valid metric and reached opposite conclusions. You choose which fairness you value most and defend that choice openly rather than searching for a metric that satisfies everything.
How do you measure fairness gaps in practice?
Use open-source libraries rather than hand-rolling metrics: Fairlearn provides demographic_parity_difference and equalized_odds_difference, and IBM's AIF360 offers a broader toolkit of metrics and mitigation algorithms. Both Fairlearn functions return a gap where zero means perfect fairness. Run them on a held-out test set, never on training data, report confidence intervals for small groups, and check intersections — a model can look fair for gender and race separately while failing for specific combinations.
Wrapping This Up
Demographic parity compares outcomes, asking whether groups receive positive predictions at equal rates. Equalized odds compares errors, asking whether the model treats equally qualified people equally. Neither wins everywhere, and the math proves you can't satisfy every definition at once.
Choose based on your use case, your label quality, and the harm each error causes. Then measure, document, and revisit the choice as your data changes.
Here's your next step: run both metrics on your current model this week and look at the gap. If the numbers surprise you, good. Surprise means you learned something before your users did.


