Contents
Figure 1: The denial letter says why. The applicant wants to know what to do next
A loan applicant is denied. A feature attribution tells them that income and debt ratio drove the decision, which is true and almost useless. What they actually want to know is what to do about it. A counterfactual explanation answers that: "You would have been approved if your debt ratio were 0.31 or lower and your credit history were six years or longer." It describes the smallest change to the input that flips the model's decision, which is why counterfactuals are often the most actionable form of explanation.
What a Counterfactual Is
Counterfactual explanations are post-hoc techniques that identify the smallest input changes needed to alter a prediction. They don't open up the model. They probe it from the outside and report a nearby input that would have produced a different outcome, which makes them model-agnostic and easy for non-technical audiences to understand.
A good counterfactual satisfies several properties, and the literature is fairly consistent on the list:
- Validity: the changed input actually gets the desired prediction.
- Proximity: it stays close to the original instance.
- Sparsity: it changes as few features as possible.
- Plausibility: it stays near realistic data, not impossible combinations.
- Actionability: it changes only features a person can actually change.
- Causality: it respects real-world cause-and-effect constraints.
- Diversity: it offers several distinct routes, not just one.
Counterfactual vs Recourse
A subtle distinction matters here. A counterfactual describes a nearby input with a different outcome. Algorithmic recourse adds a requirement: the changes must be something the person can realistically do. Researchers noted that plain counterfactuals ignore the actionability of the prescribed changes, which recourse is designed to address. A counterfactual saying "be 10 years younger" is mathematically valid and practically absurd, and a recommended counterfactual should never change immutable features.
A Hands-On Example with DiCE
DiCE (Diverse Counterfactual Explanations) is one of the best-known libraries. It was designed to balance diversity, sparsity, and proximity using a combined loss, and it supports several search strategies including random sampling, genetic search, and gradient-based optimization depending on the model backend. Install it with pip install dice-ml. Here's a self-contained example using synthetic loan data, so no downloads are needed:
import numpy as np
import pandas as pd
import dice_ml
from sklearn.ensemble import RandomForestClassifier
rng = np.random.default_rng(0)
n = 4000
df = pd.DataFrame({
"income": rng.normal(55_000, 15_000, n).clip(15_000),
"debt_ratio": rng.uniform(0.05, 0.8, n),
"credit_years": rng.integers(0, 25, n),
"age": rng.integers(21, 70, n),
})
score = (df.income / 20_000) - 4 * df.debt_ratio + 0.08 * df.credit_years \
+ rng.normal(0, 0.4, n)
df["approved"] = (score > 1.5).astype(int)
X = df.drop(columns="approved")
clf = RandomForestClassifier(n_estimators=200, random_state=0).fit(X, df["approved"])
data = dice_ml.Data(
dataframe=df,
continuous_features=["income", "debt_ratio", "credit_years", "age"],
outcome_name="approved",
)
model = dice_ml.Model(model=clf, backend="sklearn")
explainer = dice_ml.Dice(data, model, method="random")
query = X[df["approved"] == 0].iloc[[0]] # a denied applicant
cfs = explainer.generate_counterfactuals(
query,
total_CFs=3,
desired_class="opposite",
features_to_vary=["income", "debt_ratio", "credit_years"], # age stays fixed
permitted_range={"debt_ratio": [0.05, 0.6], "credit_years": [0, 25]},
)
cfs.visualize_as_dataframe(show_only_changes=True)
A few points deserve attention. features_to_vary is how you encode actionability: age is excluded because the applicant can't change it. permitted_range keeps suggestions inside realistic bounds, so the library won't suggest a negative debt ratio. desired_class="opposite" flips the decision, and total_CFs=3 asks for three different routes. DiCE's API evolves, so check the current documentation if an argument errors.
Always verify validity yourself instead of trusting the output blindly:
cf_df = cfs.cf_examples_list[0].final_cfs_df
print(clf.predict(cf_df[X.columns])) # should be all 1s
Reading and Presenting the Results
DiCE returns several alternatives, such as one that lowers debt ratio, another that raises income and credit history together. Present these as options, not instructions: "Any one of these changes would flip the decision." Translate the numbers into plain language, and make sure each option names only features the person controls. Prefer two or three diverse, sparse options over a long list, since an explanation nobody can act on isn't helping anyone.
The Methods Landscape
Counterfactual generation has grown into a crowded field. Early methods like Wachter's approach treated it as an optimization problem, finding the closest point across the decision boundary. DiCE added diversity and flexible search. Other methods address plausibility more directly: FACE uses density estimates to find paths through realistic data, though that approach scales poorly to high-dimensional data and handles categorical variables inadequately. Generative-model approaches like C-CHVAE keep counterfactuals on the data manifold by working in a learned latent space. Newer work such as LiCE models feature domains and uses likelihood to enforce plausibility, and several recent frameworks add causal constraints or domain-expert rules so suggestions respect real-world relationships.
DiCE itself has known gaps. It emphasizes diversity but does not enforce plausibility through a data distribution, and researchers have observed it can lack robustness, producing explanations that are sensitive to small perturbations. Extensions such as DiCE-Extended target that weakness directly. Alibi, mentioned in the tools piece in this series, offers its own counterfactual methods too, subject to its license terms.
The Properties Conflict
Here's the uncomfortable part: the desirable properties pull against each other. High diversity can come at the cost of staying close to the training data, and high validity can come at the cost of proximity. A counterfactual that changes one feature barely might be implausible, while one that's perfectly plausible might require changing five. There's no single best counterfactual, only a tradeoff you have to choose on purpose, guided by what the audience can act on.
Where Counterfactuals Mislead
- They're statistical, not causal. A counterfactual describes what the model would predict, not what would happen in the real world. If "increase income" correlates with other features that the model uses, changing income alone may not produce the outcome in reality. Analyses summarized in the literature report that off-the-shelf machine learning counterfactuals can conflict with true causal counterfactuals, in some common causal structures at notable rates, so recourse advice can fail to achieve its intended real-world effect.
- They can be fragile. Small changes to the input or the model can invalidate a counterfactual, so an applicant who follows advice precisely may still be denied. Test robustness by perturbing counterfactuals slightly and re-checking validity.
- They expire when the model changes. Retraining shifts the decision boundary, so a counterfactual that was valid last quarter may not hold. Communicate that explanations apply to the model version that produced them.
- Actionability is context-dependent. Whether "reduce debt ratio by 0.2" is realistic depends on the person, so involve domain experts and avoid suggestions that quietly assume resources people don't have.
- Fairness cuts both ways. The cost of achieving recourse can differ across groups. If one group needs much larger changes to flip a decision, that's evidence of disparity worth auditing, not just an individual explanation.
- They reveal the decision boundary. Handing out many counterfactuals through an API gives outsiders information about how your model decides, which can enable gaming or model extraction. Rate-limit or aggregate where that's a concern.
Counterfactuals and Regulation
Counterfactuals fit regulatory goals well because they offer meaningful, individual-level explanations without exposing model internals. That's why they're often paired with feature attributions for compliance, a pattern mentioned in the XAI beginner's guide and illustrated by the SHAP values walkthrough. In credit contexts, adverse action notices already require telling people why they were denied, and counterfactuals can make that notice more useful. I'm not a lawyer, though, and what counts as a sufficient explanation varies by jurisdiction and use case, so check with legal and compliance before relying on counterfactuals as your compliance mechanism.
A Practical Workflow
- Define which features are mutable, which are immutable, and which are only conditionally actionable, with domain input.
- Set realistic ranges and constraints before generating anything.
- Generate several diverse candidates, then filter for validity, sparsity, and plausibility.
- Verify each by re-running the model, and test small perturbations for robustness.
- Review a sample with domain experts and, where possible, with people affected by the decision.
- Present two or three options in plain language with the model version and date noted.
- Log the counterfactuals alongside predictions for audit.
Common Pitfalls
- Skipping actionability constraints and presenting suggestions that change age, location history, or other things people can't change.
- Treating counterfactuals as causal advice, promising outcomes the real world won't deliver.
- Failing to check validity after generation, especially with stochastic search methods.
- Ignoring plausibility, so suggestions land in regions with no real training data.
- Presenting a single counterfactual as the answer when several distinct routes exist.
- Forgetting that model updates invalidate earlier explanations.
Recommended Books
| Cover | Book | Description | Get it |
|---|---|---|---|
![]() |
Interpretable Machine Learning | covers counterfactual methods alongside SHAP and LIME, with the property tradeoffs spelled out. | View on Amazon |
![]() |
The Book of Why | the causal reasoning background behind knowing when a counterfactual is more than a statistical description. | View on Amazon |
![]() |
Interpretable Machine Learning with Python | hands-on companion matching the DiCE code above. | View on Amazon |
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What is a counterfactual explanation?
It is the smallest change to an input that flips a model's decision, reported as a nearby instance with a different outcome — for example, "approved if debt ratio were 0.31 or lower and credit history were six years or longer." Counterfactuals are post-hoc and model-agnostic: they probe the model from the outside without opening it up, which makes them easy for non-technical audiences to understand and directly actionable compared with feature attributions.
What is the difference between a counterfactual and algorithmic recourse?
A counterfactual describes a nearby input with a different outcome; algorithmic recourse adds the requirement that the changes be something the person can realistically do. Plain counterfactuals ignore the actionability of the prescribed changes, so "be 10 years younger" can be mathematically valid and practically absurd. In practice, encode actionability with features_to_vary, permitted ranges, and immutable-feature exclusions, and never change age or other attributes a person cannot control.
How do you make counterfactuals realistic and trustworthy?
Set mutable versus immutable features with domain input before generating anything, constrain ranges to realistic values, and request several diverse candidates rather than one. Then verify each by re-running the model, perturb it slightly to test robustness, and check that it lands in plausible data regions rather than impossible feature combinations. Generate, filter, verify, and review with domain experts — never present raw generator output as advice.
Are counterfactual explanations causal?
No. A counterfactual describes what the model would predict, not what would happen in the real world, and analyses report that off-the-shelf machine learning counterfactuals can conflict with true causal counterfactuals at notable rates in common causal structures. Treat them as statistical descriptions of the decision boundary: pair them with domain knowledge, state which model version produced them, and involve legal and compliance teams before treating them as a compliance mechanism.
Wrapping This Up
Counterfactual explanations answer the question people actually ask, what would change this decision, by finding a nearby input that flips the prediction. Libraries like DiCE make them easy to generate with constraints for actionability, diversity, and plausible ranges, and the output works well alongside SHAP-style attributions for a complete picture. The catch is that counterfactuals describe the model, not reality, so causal validity, robustness, and fairness need deliberate checking.
Will a counterfactual tell someone exactly how to get approved? Not with certainty, because it reflects what a model predicts rather than what the world guarantees. But with sensible constraints, verification, and honest framing, it turns "the model decided" into "here are concrete changes that would have made a difference," which is the explanation people were asking for in the first place.


