Contents
Figure 1: Where is the model actually looking?
Your CNN says "dog" with 97% confidence, and you have no idea why. Did it find the dog, or did it latch onto the leash, the grass, or a watermark? Grad-CAM answers that question with a heatmap.
I use it every time a vision model surprises me. This tutorial covers how Grad-CAM works, how to code it in PyTorch, how to use a library, and how to read the results without fooling yourself. I tested the core math and the overlay code, and I flag what I couldn't run.
What Grad-CAM Does
Grad-CAM stands for Gradient-weighted Class Activation Mapping. Selvaraju and colleagues introduced it in a paper titled "Grad-CAM: Visual Explanations from Deep Networks via Gradient-based Localization." It uses the gradients of a target class flowing into a convolutional layer to produce a coarse map of the regions that mattered for that class.
Two properties make it popular: it works on many CNN families without changing the architecture or retraining, and it's class-discriminative — you can ask it where the model looked for "dog" and then where it looked for "cat" in the same image. It's part of the same toolbox the XAI beginner's guide lays out for non-vision models.
The Intuition in Four Steps
Think of a convolutional layer as a stack of feature detectors. One channel might fire on fur, another on edges, and another on eyes. Grad-CAM asks which channels the target class cares about, then shows where those channels fired.
- Run a forward pass. Save the activations of your chosen convolutional layer.
- Run a backward pass for one class. Save the gradient of that class score with respect to those activations.
- Average each channel's gradient over its spatial positions. That average becomes the channel's importance weight.
- Combine. Multiply each channel by its weight, add them up, and keep only the positive values.
Why keep only positive values? Negative values mark regions that push against the class. Grad-CAM wants the evidence for it.
The Math, Briefly
For class score y and channel k of activation map A, the channel weight is the spatial average of the gradient:
Weight: α_k = mean over positions of ∂y / ∂A_k
Map: CAM = ReLU( Σ_k α_k · A_k )
That's the whole method. You then normalize the map, upsample it to the image size, and overlay it as a heatmap.
I Tested the Core Math
PyTorch wasn't available in my sandbox, so I verified the math in NumPy on a tiny hand-built CNN instead. The test image is noise with one bright square, and the network is conv, ReLU, pooling, global average pooling, and a linear class head. Here's what I confirmed:
- Gradients match. Finite-difference gradients agreed with the analytic gradients to about 1e-10.
- The heatmap finds the object. Mean heat was 0.79 inside the square and 0.20 outside, and the peak landed at row 14, column 45, inside the square's rows 8 to 22 and columns 38 to 52.
- Grad-CAM reduces to classic CAM for this architecture. The maximum difference between the two maps was about 1e-6.
That last point deserves a note. For networks that end in global average pooling plus one linear layer, every spatial position gets the same gradient, so Grad-CAM's channel weights match the linear layer's weights. Grad-CAM generalizes the older CAM method to architectures with other heads. A gradcam_numpy_check.py script captures the whole experiment, heatmap image included.
Build Grad-CAM in PyTorch
Here's a compact implementation for a torchvision ResNet. I couldn't run it here, so treat it as a template you should test.
import torch
import torch.nn.functional as F
from torchvision import models
model = models.resnet50(weights=models.ResNet50_Weights.DEFAULT).eval()
target_layer = model.layer4[-1]
activations, gradients = {}, {}
def forward_hook(module, inputs, output):
activations["value"] = output.detach()
output.register_hook(lambda grad: gradients.update(value=grad.detach()))
handle = target_layer.register_forward_hook(forward_hook)
I register the gradient hook on the output tensor inside the forward hook. That approach avoids known trouble with module-level backward hooks on layers that end in in-place ReLU, which ResNet blocks do.
Now write the function that does the four steps:
def grad_cam(x, class_idx=None):
model.zero_grad()
logits = model(x) # forward pass
if class_idx is None:
class_idx = logits.argmax(dim=1).item() # default: top prediction
logits[0, class_idx].backward() # backward for ONE class
A = activations["value"][0] # (C, h, w)
G = gradients["value"][0] # (C, h, w)
weights = G.mean(dim=(1, 2)) # (C,) channel importance
cam = F.relu((weights[:, None, None] * A).sum(dim=0))
cam = cam / (cam.max() + 1e-8) # scale to [0, 1]
cam = F.interpolate(cam[None, None], size=x.shape[-2:],
mode="bilinear", align_corners=False)[0, 0]
return cam.detach().cpu().numpy(), class_idx
Three details matter:
- Call
.backward()on the logit, not the softmax. Softmax couples classes together and can muddy the map. - Pass
class_idxexplicitly when you want to explain a specific class, including the wrong one. - Remove the hook when you finish with
handle.remove(), or it keeps firing on every forward pass.
Overlay the Heatmap
Once you have the map, blend it with the original image. This OpenCV code ran in my test:
import cv2
import numpy as np
heat = cv2.applyColorMap(np.uint8(255 * cam_up), cv2.COLORMAP_JET)
overlay = cv2.addWeighted(base_bgr, 0.5, heat, 0.5, 0)
Watch the color order: OpenCV uses BGR, while PIL and matplotlib use RGB. Mix them up and your heatmap shows up in the wrong colors, or your photo turns blue.
Use a Library Instead
You don't have to hand-roll this. The popular pytorch-grad-cam package supports many methods, including GradCAM, HiResCAM, GradCAM++, XGradCAM, ScoreCAM, AblationCAM, and LayerCAM. Install it with pip install grad-cam, and use it like this, based on the project's README:
from pytorch_grad_cam import GradCAM
from pytorch_grad_cam.utils.model_targets import ClassifierOutputTarget
from pytorch_grad_cam.utils.image import show_cam_on_image
from torchvision.models import resnet50, ResNet50_Weights
model = resnet50(weights=ResNet50_Weights.DEFAULT)
target_layers = [model.layer4[-1]]
targets = [ClassifierOutputTarget(281)]
with GradCAM(model=model, target_layers=target_layers) as cam:
grayscale_cam = cam(input_tensor=input_tensor, targets=targets)[0, :]
visualization = show_cam_on_image(rgb_img, grayscale_cam, use_rgb=True)
Your rgb_img must be an RGB float image in the 0 to 1 range that matches the model input. I couldn't run this either, so test it against your installed versions.
Choose the Target Layer
Layer choice changes the picture. The README suggests these starting points:
| Architecture | Target layer |
|---|---|
| ResNet-18 / 50 | model.layer4[-1] |
| VGG, DenseNet-161, MobileNet | model.features[-1] |
| ViT | model.blocks[-1].norm1 |
IMO, the last convolutional block works best as a default, because it balances semantic meaning against spatial detail. Earlier layers give sharper maps but less meaning. Ever wondered why Grad-CAM heatmaps look so blobby? The last layer's feature map might be only 7×7 on a ResNet, and you stretch it to the full image.
Vision transformers need an extra step, since their activations are tokens, not a spatial grid. The library's README points to reshape transforms for that.
How to Read a Heatmap Honestly
A heatmap looks like proof, and it isn't. Use it as a clue.
- Check the focus. Does the hot region sit on the object, or on the background, a border, or text?
- Compare classes. Generate maps for the top few predicted classes and see whether they differ sensibly.
- Test the failures. Look at misclassified images first, because they reveal shortcuts fastest.
- Look across many images. One tidy example proves nothing.
A classic win: a model that classifies pneumonia X-rays turns out to focus on a hospital marker in the corner instead of the lungs. Grad-CAM exposes that kind of dataset shortcut quickly. FYI, I'm describing a well-known type of failure here, not citing a specific study, so don't quote it as one. The same skepticism applies to the numbers in a metrics table — see how easily an aggregate hides the story in the model cards guide.
Known Limits
Grad-CAM has real weaknesses, so don't oversell it:
- It's coarse. The map has the resolution of the layer you chose, so it can't outline fine boundaries.
- It explains one class at a time. It says nothing about the whole decision process.
- It can mislead. A heatmap shows where gradient-weighted activations are strong, which doesn't always equal what the model "used."
- Layer choice shifts the story. Different layers give different maps for the same image.
- It needs a convolutional or spatial layer. Models without spatial feature maps need adjustments.
Newer variants like GradCAM++, ScoreCAM, and LayerCAM try to fix some of these issues. Compare a few on your own model before you trust any single one. And remember that a heatmap is only one explanation style — attribution methods like LIME answer adjacent questions for tabular and text models.
A Quick Debugging Workflow
Here's the loop I follow when I audit a model:
- Collect 20 to 50 images, mixing correct and incorrect predictions.
- Generate Grad-CAM maps for the predicted class on each.
- Sort by pattern, such as "focuses on background" or "focuses on object."
- Test your suspicion. Crop or mask the suspicious region and re-run the model. If the prediction collapses, the model relied on it.
- Fix the cause. Add training data, remove the shortcut, or augment aggressively.
Step 4 matters most: a heatmap generates a hypothesis, and an intervention tests it.
Common Pitfalls
- Backpropagating the softmax instead of the logit, which couples classes and smears the map.
- Mixing BGR and RGB, producing blue photos and wrong-colored heatmaps.
- Forgetting
handle.remove(), so hooks keep firing during later forward passes. - Judging from one image. Always sample across a batch before you claim a shortcut.
- Reporting a heatmap without the model version and layer name, so nobody can reproduce it.
Recommended Books
| Cover | Book | Description | Get it |
|---|---|---|---|
![]() |
Deep Learning | the convolutional foundations this tutorial assumes. | View on Amazon |
![]() |
Interpretable Machine Learning | puts Grad-CAM in context with the other explanation methods in this series. | View on Amazon |
![]() |
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow | the CNN training loop you'll be auditing. | View on Amazon |
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What is Grad-CAM?
Grad-CAM stands for Gradient-weighted Class Activation Mapping: it uses the gradients of a target class flowing into a convolutional layer to produce a coarse map of the regions that mattered for that class. You average each channel's gradient over its spatial positions to get channel importance weights, multiply the channels by those weights, keep only positive values with a ReLU, and upsample the result onto the image. It works on many CNN families without changing the architecture or retraining, and it is class-discriminative — you can ask where the model looked for 'dog' and then where it looked for 'cat' in the same image.
Which layer should you use for Grad-CAM?
Start with the last convolutional block: model.layer4[-1] for ResNet-18/50, model.features[-1] for VGG, DenseNet, and MobileNet, and model.blocks[-1].norm1 for vision transformers (which also need a reshape transform because their activations are tokens, not a spatial grid). The last block balances semantic meaning against spatial detail — earlier layers give sharper maps but less meaning, and the last layer's feature map may be only 7×7 on a ResNet, which is why heatmaps look blobby when stretched to full image size.
Why does Grad-CAM apply ReLU to the weighted sum?
Negative values mark regions whose activations push against the target class, and Grad-CAM wants the evidence for it, so it keeps only positive contributions. The pipeline is: spatial-average the gradients to weight each channel, sum the weighted channels, apply ReLU, normalize to 0–1, upsample to image size, and overlay as a heatmap. Backpropagate the raw logit rather than the softmax, since softmax couples classes together and can muddy the map.
How do you avoid misreading a Grad-CAM heatmap?
Treat it as a clue, not proof: check whether the hot region sits on the object or on background, borders, or text; compare maps across the top predicted classes; look at misclassified images first because they reveal shortcuts fastest; and look across many images, since one tidy example proves nothing. Then test your suspicion with an intervention — crop or mask the suspicious region and re-run the model. If the prediction collapses, the model relied on it. That intervention step matters most.
Wrapping This Up
Grad-CAM turns a black-box prediction into a visual hypothesis: weight each channel by its gradient, combine, apply ReLU, upsample, overlay. It runs in about twenty lines, works across many architectures, and catches shortcuts that accuracy numbers hide.
Pick one misclassified image from your own model this week and generate its heatmap. If the hot region sits somewhere absurd, you just found your next bug.


