Sam Austin on October 8, 2026

Explainable AI (XAI) for Beginners: Complete Guide

Explainable AI (XAI) for Beginners: Complete Guide
Contents

Loan application paperwork and a calculator on a desk

Figure 1: "The model decided" is not an answer a bank can give anymore

A bank denies your loan application. You ask why, and the answer is "the model decided." That response is increasingly unacceptable to customers, regulators, and the engineers who have to debug the model. Explainable AI (XAI) is the set of techniques for answering "why did the model do that?" in terms a human can check. This guide covers the vocabulary, the main methods, how to read their output, and where they mislead.

Interpretability vs Explainability

These terms get used interchangeably, but there's a useful distinction. Interpretability means the model's inner workings are understandable by design. A short decision tree or a linear regression qualifies, since you can read the rules or coefficients directly. Explainability means generating a human-understandable account of an opaque model's behavior after the fact. A gradient-boosted ensemble or a neural network isn't readable, but techniques like SHAP or LIME can still explain individual predictions.

Put simply: explainability focuses on justifying individual predictions in terms users and stakeholders can act on, while interpretability seeks a general understanding of internal mechanisms.

Why It Matters

Four motivations come up repeatedly.

  • Debugging: explanations reveal when a model relies on a spurious feature, like a hospital ID that correlates with outcomes.
  • Trust: domain experts accept models more readily when the reasoning matches their knowledge.
  • Fairness: seeing which features drive decisions helps surface proxy discrimination — the same concern the bias in synthetic data article raises about generators hiding group effects.
  • Regulation: the EU AI Act's Article 13 requires high-risk systems to be transparent enough for deployers to interpret outputs, though it doesn't prescribe specific methods like SHAP or LIME. In the US, adverse action notice rules for credit decisions create similar practical pressure.

Two Routes: Interpretable by Design or Explained After the Fact

You have two broad strategies. The first is using an inherently interpretable model — linear models, small trees, or rule lists. The second is training a more powerful black-box model and applying post-hoc explanation tools.

Researchers like Cynthia Rudin have argued that for high-stakes decisions, you should prefer interpretable models over explaining black boxes, since an explanation is only an approximation of what the model really does. That's a serious point worth weighing: if a simple model performs nearly as well on your problem, the simpler model removes the whole explanation-fidelity question — the same "skip the baseline and you won't know" logic from the small datasets ladder.

Global vs Local Explanations

A global explanation describes the model's overall behavior: which features matter most across all predictions. A local explanation covers one prediction: why this applicant was denied. You usually need both. Global views help you validate and debug the model, while local views answer individual "why" questions.

The Beginner's Toolbox

Permutation importance is the simplest starting point. Shuffle one feature's values in a validation set and measure how much model performance drops. A big drop means the model leans on that feature. It's global only, but easy to understand and hard to misuse.

SHAP (SHapley Additive exPlanations) uses game theory to attribute a prediction to each feature, giving both local explanations and global ones when you aggregate values across a dataset. SHAP and LIME remain the workhorses for tabular and classical machine learning. SHAP's guarantees around consistency make it easier to defend, and its results are deterministic, which matters when you need the same answer every time. Its main drawback is computational cost on large models or datasets — the SHAP values article walks through the mechanics on a worked example.

LIME (Local Interpretable Model-agnostic Explanations) fits a simple model, usually linear, around one prediction by perturbing the input and observing outputs. It's fast and intuitive, but stochastic: running it twice on the same instance can produce different explanations, which is a real problem when consistency is expected. Increasing the number of perturbation samples reduces that variance.

Counterfactual explanations answer "what would need to change for a different outcome?" For example, "approved if income were $4,000 higher." They don't explain how the model works internally, only what small input changes would flip the result — often more actionable and, for consumer-facing explanations, more meaningful than feature attributions. Libraries like DiCE and Alibi generate them.

Attribution for deep networks uses methods like Integrated Gradients and saliency maps, available in libraries such as Captum. These highlight which pixels or tokens influenced a prediction. Treat them cautiously: research on saliency map sanity checks showed some methods produce plausible-looking maps even when the model's weights are randomized, so verify that an attribution method actually responds to the model before trusting it.

A Quick SHAP Example

This example trains a model on scikit-learn's built-in diabetes dataset and produces both a global and a local view:

import shap
from sklearn.datasets import load_diabetes
from sklearn.ensemble import GradientBoostingRegressor
from sklearn.model_selection import train_test_split

X, y = load_diabetes(return_X_y=True, as_frame=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0)

model = GradientBoostingRegressor(random_state=0).fit(X_train, y_train)

explainer = shap.TreeExplainer(model)
shap_values = explainer(X_test)

shap.plots.beeswarm(shap_values)       # global: which features matter overall
shap.plots.waterfall(shap_values[0])   # local: why this one prediction

Install it with pip install shap. The beeswarm plot ranks features by overall impact and shows whether high or low values push predictions up or down. The waterfall plot walks through one prediction, starting from the average output and adding each feature's contribution. Plot APIs shift between SHAP releases, so check the current docs if a call errors.

SHAP or LIME?

Use SHAP when you need consistent, defensible explanations, global summaries, or compliance documentation. Use LIME for quick, intuitive debugging during development, when speed matters more than stability. Many teams use both: LIME while exploring, SHAP for anything that has to hold up later. For regulated or consumer-facing use, pairing SHAP-style attribution with counterfactuals gives you both the technical account and the user-friendly one.

Explaining LLMs Is a Different Problem

Tabular techniques don't scale cleanly to language models. LIME and SHAP become computationally impractical at billions of parameters, and the LLM-specific alternatives have their own flaw: chain-of-thought reasoning and post-hoc citations often prioritize plausibility over faithfulness.

Faithfulness means the stated reasoning actually caused the answer. A model can write a convincing explanation that has nothing to do with how it reached its output — so LLM self-explanations may be plausible-sounding confabulations rather than accurate accounts of internal processing. Practical responses include perturbation and counterfactual tests (change the stated reasoning and check whether the answer changes), grading explanations with a separate faithfulness evaluator, and mechanistic interpretability, which studies model internals directly but is mostly still research. Treat an LLM's explanation of itself as a claim to test, not evidence.

How to Read Explanations Without Fooling Yourself

A few traps catch nearly every beginner.

  • An explanation describes the model, not the world. If SHAP says income drives predictions, that tells you what the model uses, not that income causes the outcome. Don't read attributions as causal claims.
  • Correlated features share credit unpredictably. When two features carry similar information, attributions can split between them in ways that mislead, so inspect correlations before drawing conclusions.
  • Plausible isn't the same as faithful. A neat explanation can still misrepresent the model. For high-stakes decisions, have domain experts review explanations and treat them as evidence, not ground truth.
  • Instability is a signal. If explanations change drastically with tiny input changes or random seeds, that tells you something about the explanation method, the model, or both.

Regulation in Brief

The EU AI Act sets transparency expectations for high-risk systems, and GDPR contains provisions around automated decision-making. Implementation timelines for parts of the AI Act have been phasing in through 2026 and 2027, and exact dates and standards are still evolving, so check current status for your system rather than relying on a fixed date.

Two practical rules cut through the detail. Regulations generally require meaningful explanations rather than a particular technique. Store the explanations you generate, since being able to produce them later matters as much as generating them at all — the same audit discipline production model monitoring applies to metrics. And as always, consult your legal or compliance team about what your specific use case requires.

A Practical Workflow for Beginners

  1. Start by asking whether an interpretable model is good enough, and compare it against your black-box candidate on your actual metric.
  2. If you use a complex model, begin with permutation importance for a sanity check on what it relies on.
  3. Add SHAP for global and local views, and examine the top features for anything suspicious or unethical.
  4. Add counterfactuals if people affected by decisions need actionable explanations.
  5. Check stability by rerunning explanations with different seeds and slightly perturbed inputs.
  6. Have a domain expert review a sample of explanations, and store explanations alongside predictions for auditability.

Common Pitfalls

  • Treating feature attributions as causal explanations, which overstates what they show.
  • Presenting LIME output from a single run, without checking whether it stays stable.
  • Trusting LLM chain-of-thought as a faithful account of the model's reasoning.
  • Explaining a model that isn't any better than a simple interpretable alternative.
  • Skipping validation of the explanation method itself, especially saliency-style methods.
  • Generating explanations only for demos and never storing them where auditors or users can retrieve them.
CoverBookDescriptionGet it
Cover of “Interpretable Machine Learning” Interpretable Machine Learningby Christoph Molnar the free book this space is built on: methods, limits, and the model-agnostic toolkit explained properly. View on Amazon
Cover of “Explainable AI: Interpreting, Explaining and Visualizing Deep Learning” Explainable AI: Interpreting, Explaining and Visualizing Deep Learning edited by Samek, Montavon & Müller — deeper coverage of attribution methods for networks, including the sanity checks discussed here. View on Amazon
Cover of “Interpretable Machine Learning with Python” Interpretable Machine Learning with Python hands-on companion matching the SHAP and counterfactual code above. View on Amazon

Unlock AI That Actually Works

Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.

Click here to get GPTAstra Max now — one-time payment, lifetime access.

Frequently Asked Questions

What is the difference between interpretability and explainability?

Interpretability means the model's inner workings are understandable by design — a short decision tree or linear regression you can read directly. Explainability means generating a human-understandable account of an opaque model's behavior after the fact, with techniques like SHAP or LIME. Interpretability seeks a general understanding of internal mechanisms; explainability justifies individual predictions in terms users and stakeholders can act on.

Should I use SHAP or LIME?

Use SHAP when you need consistent, defensible explanations, global summaries, or compliance documentation — its consistency guarantees and deterministic results make it easier to defend, with computational cost as the main drawback. Use LIME for quick, intuitive debugging during development, when speed matters more than stability: it's stochastic, so two runs on the same instance can produce different explanations (increase perturbation samples to reduce that variance). Many teams use both: LIME while exploring, SHAP for anything that has to hold up later.

What are counterfactual explanations?

They answer "what would need to change for a different outcome?" — for example, "approved if income were $4,000 higher." They don't explain how the model works internally, only what small input changes would flip the result, which is often more actionable and, for consumer-facing explanations, more meaningful than feature attributions. Libraries like DiCE and Alibi generate them.

Are LLM chain-of-thought explanations reliable?

Treat them as claims to test, not evidence. Faithfulness means the stated reasoning actually caused the answer, and a model can write a convincing explanation that has nothing to do with how it reached its output — plausible-sounding confabulation rather than an accurate account of internal processing. Test them with perturbation and counterfactual checks (change the stated reasoning and see whether the answer changes) and grade explanations with a separate faithfulness evaluator; mechanistic interpretability is promising but mostly still research.

Wrapping This Up

XAI gives you tools for asking why a model behaved the way it did: permutation importance for a quick global check, SHAP for consistent local and global attributions, LIME for fast exploratory insight, counterfactuals for actionable "what would change this" answers, and specialized methods for deep networks and LLMs. Each carries limits, from SHAP's cost and LIME's instability to the faithfulness problem that makes LLM self-explanations unreliable.

Will an explanation tell you what a model truly does? Not perfectly, since explanations are approximations that need validation. But used carefully — with stability checks, expert review, and a willingness to prefer simpler models when they're good enough — they turn a black box into something you can question, debug, and defend.

What are You Looking For?

esc