Contents
Figure 1: Which token is looking at which
You trained a transformer, and it works. Then someone asks why it predicted "positive" for a sentence that sounds neutral, and you shrug. Two tools help you answer: attention visualization, which shows where the model looks, and probing, which tests what its hidden states know.
Both tools mislead people constantly. This guide shows you how to use each one and how to avoid the traps. I tested the core code on toy data and report real numbers below.
Two Questions, Two Tools
Interpretability starts with a choice of question. Attention visualization asks, "Which tokens did this layer mix together?" Probing asks, "Does this hidden state contain information about X?"
Those questions sound close, but they're not the same: attention describes routing, and probing describes content. Keep that distinction in mind and you'll avoid most beginner mistakes.
Attention Basics
Self-attention lets every token build its new representation from a weighted mix of other tokens. Each token produces a query, a key, and a value. The model scores every query against every key, applies a softmax, and uses the resulting weights to blend the value vectors.
That gives you a matrix per head: rows are the tokens doing the looking, and columns are the tokens getting looked at. Each row sums to 1. I verified that property in a toy NumPy transformer, and I also confirmed that a causal mask drives every weight above the diagonal to zero.
Read the Heatmap
A heatmap of that matrix is the classic attention plot. Read it row by row: pick a token on the left and see where its weight goes on the right. A saved attention_heatmap.png from my toy sentence "the cat sat on the mat" shows exactly the patterns below.
Watch for these patterns:
- Diagonal stripes: tokens attend mostly to themselves.
- Previous-token stripes: tokens attend to their left neighbor.
- Vertical bars: many tokens attend to one special token, such as
[CLS]or a separator. - Diffuse blur: the head spreads weight almost evenly, so it tells you little.
Ever wondered how "flat" a head is? Compute its entropy per row. In my toy run, entropy ranged from 1.23 to 1.76, against a uniform ceiling of 1.79, so the toy heads stayed fairly diffuse.
Pull Attention From a Real Model
With Hugging Face Transformers, you request attention weights when you load the model. This pattern comes from the BertViz README:
from transformers import AutoTokenizer, AutoModel
model_name = "bert-base-uncased"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModel.from_pretrained(model_name, output_attentions=True)
inputs = tokenizer.encode("The cat sat on the mat", return_tensors="pt")
outputs = model(inputs)
attention = outputs[-1] # one tensor per layer
tokens = tokenizer.convert_ids_to_tokens(inputs[0])
Each layer's tensor has the shape (batch, heads, query tokens, key tokens). I couldn't run this snippet here, because PyTorch wasn't available in my sandbox. Also note that newer library versions may use faster attention kernels that don't return weights, so check the docs if attention comes back empty.
Visualize With BertViz
BertViz turns those tensors into interactive plots. Install it with pip install bertviz, then call it in a Jupyter notebook:
from bertviz import head_view, model_view
head_view(attention, tokens)
model_view(attention, tokens)
It offers three views:
- Head view: shows attention for one or more heads in the same layer.
- Model view: gives a bird's-eye view across all layers and heads.
- Neuron view: shows the query and key neurons behind each weight, but it requires BertViz's own BERT, GPT-2, or RoBERTa classes, not standard Hugging Face models.
On long inputs, pass include_layers=[5, 6] to limit the load. Large inputs can crash a Colab runtime.
Why Attention Isn't a Full Explanation
Here's where people go wrong: a bright cell feels like proof that "the model used this token." The research disagrees.
Jain and Wallace published "Attention is not Explanation" in 2019. They found that attention weights often don't correlate with gradient-based importance measures, and that very different attention distributions can produce equivalent predictions. Wiegreffe and Pinter replied with "Attention is not not Explanation," arguing that the answer depends on how you define explanation and proposing more rigorous tests.
IMO, both sides have a point: attention gives you a clue, but it doesn't give you a verdict. The same rule applies to the heatmaps in the Grad-CAM guide — visual evidence is a hypothesis, not a conclusion.
The Value-Vector Problem
I built a tiny demo of one reason. The attention output equals weights times value vectors, so weights alone don't determine what flows through. I gave two tokens identical value vectors, then moved attention mass between them. The output changed by about 2e-16, which is numerical zero.
In plain terms: if two tokens carry the same value, the heatmap can shift between them and the model's behavior stays put. A heatmap that ignores values misses half the computation.
The Mixing Problem
Layers also blend information. Abnar and Zuidema argue that by higher layers, information from different tokens becomes increasingly mixed, so raw attention weights stop reliably showing which input tokens matter. They propose attention rollout and attention flow to estimate contributions through the whole stack.
Rollout multiplies attention matrices across layers, usually after adding the residual connection. I implemented a simple version, with each layer's matrix set to half attention plus half identity, and confirmed the rows still sum to 1. In my toy model, the average self-weight rose from 0.146 in the last layer's raw attention to 0.222 in the rollout, because the residual path carries each token's own information forward.
Probing: Ask the Hidden States Directly
Probing takes a different route: you freeze the model, extract hidden states, and train a small classifier to predict a property from them. If a simple classifier succeeds, the hidden states contain that information.
The classic example is part-of-speech tagging. Tenney, Das, and Pavlick found in "BERT Rediscovers the Classical NLP Pipeline" that the network regions tied to each task appear in an order matching the traditional pipeline: part-of-speech tagging, then parsing, named entity recognition, semantic roles, and coreference. That result made layer-by-layer probing a standard technique.
Here's the recipe:
- Pick a property with labels, like tags, entity types, or sentiment.
- Extract hidden states from one layer, using
output_hidden_states=True. - Train a probe on those states with a held-out test set.
- Repeat for every layer to see where the information appears.
In Hugging Face, the extraction step looks like this, though I couldn't run it here:
outputs = model(**batch, output_hidden_states=True)
hidden = outputs.hidden_states # embeddings plus one tensor per layer
layer_k = hidden[k] # (batch, tokens, hidden_dim)
Then you train a scikit-learn classifier on the per-token vectors from each layer.
The Probing Trap: A Smart Probe Memorizes
A strong probe can score well without the model knowing anything. Hewitt and Liang call this out in "Designing and Interpreting Probes with Control Tasks." They propose a control task that pairs each word type with a random label. A probe can only solve it by memorizing word types, so its score measures memorization capacity.
They define selectivity as the accuracy gap between the real task and the control task. A good probe scores high on the real task and low on the control. Their experiments found that popular probes on ELMo weren't selective, and that dropout didn't fix it for MLP probes.
I reproduced the idea on synthetic data. I built word-type embeddings with and without a linearly encoded tag signal, then trained a linear probe and a 512-unit MLP probe. Chance is 0.25.
| Representation | Probe | Real task | Control task | Selectivity |
|---|---|---|---|---|
| Tag encoded | Linear | 0.977 | 0.437 | 0.541 |
| Tag encoded | MLP | 1.000 | 1.000 | 0.000 |
| No tag signal | Linear | 0.405 | 0.429 | -0.024 |
| No tag signal | MLP | 0.999 | 0.999 | 0.000 |
Read the MLP rows: on embeddings with no tag signal at all, the MLP scored 0.999 on the real task, because it memorized each word type. Its selectivity sat at zero in both cases, so it told me nothing. The linear probe separated the cases cleanly: 0.977 with the signal, 0.405 without.
The linear probe isn't perfectly clean either, since its control score of 0.437 beats chance. Check selectivity for every probe you use, and prefer simple probes.
Probing Limits
Even a selective probe has limits:
- Probes show correlation, not use. The information may sit in the hidden state without the model ever using it.
- Absence proves little. A failed probe might just be a weak probe.
- Labels carry assumptions. Your tag set shapes what you find.
- Probes depend on layer and token position. Report both.
To test whether the model uses a feature, you need interventions. Edit the representation, ablate a head, or patch activations, then watch the output change. Probing and attention suggest hypotheses, and interventions test them.
A Practical Workflow
Here's the loop I'd follow on a new model:
- Start with a behavior. Pick one concrete failure or surprise.
- Look at attention across a handful of examples for patterns, using rollout when you inspect upper layers.
- Probe layer by layer with a linear probe, and compute selectivity against a control task.
- Form a hypothesis about which layer or head matters.
- Intervene. Ablate or patch the suspect component and measure the effect.
- Report uncertainty. State what you tested and what you didn't.
FYI, steps 2 and 3 are cheap, and step 5 is where you earn your conclusions.
Common Mistakes
I've made most of these myself:
- Treating attention as a causal explanation. It's a clue about routing.
- Reading one head in one layer. Models have hundreds of heads.
- Using a big probe. High-capacity probes memorize.
- Skipping the control task. Without it, you can't read your probe's score.
- Checking a single example. One tidy heatmap proves nothing.
- Forgetting special tokens. Many heads dump weight on
[CLS]or separators, which can dominate a plot.
Recommended Books
| Cover | Book | Description | Get it |
|---|---|---|---|
![]() |
Speech and Language Processing | the classical pipeline the probing literature maps onto. | View on Amazon |
![]() |
Transformers for Natural Language Processing | attention internals with working code. | View on Amazon |
![]() |
Interpretable Machine Learning | the wider explanation-method context this series builds on. | View on Amazon |
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What is the difference between attention visualization and probing?
Attention visualization asks which tokens a given layer mixed together — it describes routing. Probing asks whether a hidden state contains information about some property — it describes content. Attention plots come from the attention weights the model already computes, while probes train a small classifier on frozen hidden states. They answer different questions, and each fails in a predictable way: attention is a clue about routing, probes show correlation rather than use, and only interventions show what the model actually uses.
Does attention explain a model's prediction?
Not on its own. Jain and Wallace's 'Attention is not Explanation' (2019) found that attention weights often don't correlate with gradient-based importance and that very different attention distributions can yield equivalent predictions; Wiegreffe and Pinter's 'Attention is not not Explanation' answered that the verdict depends on how you define explanation, proposing stricter tests. Two concrete reasons attention can mislead: identical value vectors make weight shifts invisible to the output, and information mixes across layers so raw weights in upper layers stop tracking which input tokens matter — attention rollout addresses the second problem.
What is probe selectivity and why does it matter?
Selectivity is the accuracy gap between your real task and a control task that pairs each word type with a random label — a probe can only solve the control by memorizing word types, so that score measures memorization capacity. In the toy reproduction, a 512-unit MLP scored 0.999 on a real task with no tag signal present, purely by memorizing, with selectivity at zero; a linear probe scored 0.977 with the signal and 0.405 without. High selectivity on a simple probe is the evidence; a raw probe score is not.
How do you read an attention heatmap?
Read it row by row: rows are the tokens doing the looking, columns are the tokens being looked at, and each row sums to 1. Diagonal stripes mean tokens attend mostly to themselves, previous-token stripes mean attention to the left neighbor, vertical bars mean many tokens dump weight on one special token such as [CLS], and diffuse blur means the head spreads weight almost evenly and tells you little. Entropy per row quantifies flatness, and special tokens can dominate plots if you forget to exclude them.
Wrapping This Up
Attention visualization shows routing. Probing shows content. Interventions show use. Each tool answers a different question, and each one fails in a predictable way if you over-read it. Use attention for hypotheses, linear probes with control tasks for evidence, and interventions for proof.
Pick one prediction your model gets wrong this week and look at its attention, then probe the layer where you suspect the error starts. If the story from the heatmap and the story from the probe disagree, you've found something worth investigating.


