Contents
Figure 1: How does the prediction change as the feature moves?
You've trained a gradient-boosted model, and feature importance says BMI matters. That tells you BMI is used, but not how: does risk rise steadily, level off, or jump at a threshold? Partial dependence plots (PDPs) and individual conditional expectation (ICE) plots answer that question by showing how predictions change as one feature varies. They're among the oldest and most useful model-inspection tools, and they work with any model that can make predictions.
The Core Idea
A partial dependence plot shows the average prediction as you sweep one feature across its range, with every other feature left as observed. To compute it for one feature, you:
- Pick a grid of values for the feature of interest.
- For each grid value, overwrite that feature in every row of your dataset with the grid value.
- Run the model on all those modified rows and average the predictions.
- Plot the averages against the grid.
The averaging step is the point: you're marginalizing over the other features, asking "what does the model predict on average if this feature were set to x?"
An ICE plot skips the averaging. It draws one line per data point, showing how that individual's prediction changes as the feature varies. The relationship between the two is direct: the PDP is the average of the ICE lines. In fact, scikit-learn's fast recursion method computes the PDP by implicitly averaging ICE curves, which is why that method can't produce individual lines.
Why You Want Both
The PDP summarizes, and summaries can hide things. Imagine half your customers respond to a discount with rising purchase probability and the other half respond with falling probability. The average is flat, and the PDP suggests the feature does nothing, while the model is actually reacting strongly in opposite directions. This happens whenever a feature interacts with another feature.
ICE lines expose that. If the lines fan out, cross, or split into groups, an interaction exists and the PDP's single curve is misleading. If the lines run roughly parallel, the average is a fair description of everyone. This is why scikit-learn's documentation recommends using ICE plots alongside PDPs rather than alone.
Centered ICE Plots
Raw ICE lines start at different heights, which makes it hard to compare their shapes. Centered ICE (cICE) fixes that by shifting each line to start at zero on the left edge. Centering emphasizes how individual conditional expectations diverge from the mean line, so heterogeneity in effects stands out instead of being buried in differences in baseline predictions.
A Working Example
This uses scikit-learn's bundled diabetes dataset, so no downloads are needed:
import matplotlib.pyplot as plt
from sklearn.datasets import load_diabetes
from sklearn.ensemble import HistGradientBoostingRegressor
from sklearn.inspection import PartialDependenceDisplay
from sklearn.model_selection import train_test_split
X, y = load_diabetes(return_X_y=True, as_frame=True)
X_train, X_test, y_train, y_test = train_test_split(X, y, random_state=0)
model = HistGradientBoostingRegressor(random_state=0).fit(X_train, y_train)
# 1. Classic PDP for several features
PartialDependenceDisplay.from_estimator(
model, X_train, ["bmi", "bp", "s5"], kind="average"
)
# 2. ICE lines with the PDP on top, centered
PartialDependenceDisplay.from_estimator(
model, X_train, ["bmi", "s5"],
kind="both", subsample=50, centered=True, random_state=0,
)
# 3. Two-way PDP for an interaction
PartialDependenceDisplay.from_estimator(
model, X_train, [("bmi", "s5")], kind="average"
)
plt.show()
The kind argument is the key control: "average" gives the traditional PDP, "individual" gives ICE, and "both" overlays them. subsample=50 draws 50 randomly chosen ICE lines, which keeps the plot readable and the computation fast. Note that ICE options aren't valid for two-way plots, so interaction heatmaps always use kind="average".
Fit and plot on the training data or a representative sample, since the PDP averages over whatever data you pass in. Use data that resembles what the model will see in production.
Reading the Plots
- Shape: monotonic rises, plateaus, thresholds, and U-shapes are all visible. A sudden step often marks a split the model learned on that feature.
- Scale: check the y-axis range before reading anything into a wiggle. A curve that varies by 0.2 units on a target whose spread is 100 isn't telling you much.
- Data density: the small tick marks along the x-axis show deciles of the feature's distribution. Where the ticks are sparse, the curve is built from few real examples and deserves less trust.
- Edges: by default the grid covers the 5th to 95th percentiles of each feature, which trims extreme values where the model is least reliable. You can adjust that with
percentiles, or supply specific values withcustom_values. - ICE spread: wide spread between lines means other features substantially change this feature's effect. Look for subgroups, and consider a two-way plot or a closer look at the interacting feature.
- Flat PDP, active ICE: the averaging-cancels-out case above. Never conclude a feature is irrelevant from a flat PDP alone.
Interactions with Two-Way PDPs
A two-way PDP plots the average prediction over a grid of two features at once, usually as a heatmap or contour plot. Use it when ICE lines suggest an interaction, or when you have a domain hypothesis such as "the effect of dose depends on age." Compare it against what you'd expect from the two one-way plots added together: if the surface bends or twists differently, the features interact.
Categorical Features and Classification
For categorical features, pass categorical_features so the plot treats values as categories, producing bars or point plots instead of a smooth curve. For classifiers, the plot shows the predicted probability by default, and multi-class models need a target argument to say which class to display. If you want the decision function instead of probabilities, set the response method accordingly, remembering that the scale changes, since log-odds are not bounded to the zero-to-one range.
The Big Limitation: Correlated Features
PDPs rest on an independence assumption. When you overwrite a feature with a grid value across all rows, you create combinations that may never occur in real data, for example setting age to 25 for someone whose other attributes describe a retiree. The model is then evaluated in regions where it has little or no training data, and its behavior there can be arbitrary. If the feature of interest is strongly correlated with others, the resulting curve may reflect that extrapolation, not anything real.
Practical responses:
- Check correlations first. Inspect pairwise correlations or clustering among features before trusting a PDP.
- Use density cues. Treat regions with sparse data ticks as unreliable.
- Consider ALE plots. Accumulated Local Effects, introduced by Apley and Zhu, estimate effects using only local, realistic neighborhoods of the data, avoiding the impossible-combination problem. Libraries implementing ALE exist for Python, though I haven't compared their maintenance here, so check before adopting one.
- Cross-check with other methods. Compare against SHAP dependence plots, which use different assumptions, and treat disagreement as a reason to investigate.
Other Things to Remember
- PDPs describe the model, not the world. A curve showing risk rising with a feature means the model predicts that, not that changing the feature will cause the outcome. This is the same causal caution the counterfactual explanations guide repeats: statistical descriptions of a model are not real-world guarantees.
- Cost scales with data and grid size. The brute-force method needs predictions for every row at every grid point, so a large dataset with a fine grid means many model calls. Subsample the data and lower
grid_resolutionfor exploration, since the default grid is quite fine. Recursion is faster but only works withkind="average"and certain tree-based models. - Averaging ignores sample weights unless you set them. Use the
sample_weightoption if your data needs weighting, noting it can't apply to individual ICE lines. - Interaction strength can be quantified. Friedman's H-statistic measures how much of a model's variation comes from interactions, which is a useful companion when ICE lines suggest effects vary but you want a number.
Where These Plots Earn Their Keep
- Sanity-checking models. Does the learned relationship match domain knowledge? A credit model whose predicted default risk drops as debt rises is a red flag for data problems or leakage.
- Verifying constraints. If you imposed monotonic constraints, plots confirm the model respects them.
- Finding thresholds. Sharp steps can reveal cutoffs worth investigating, such as an age at which risk jumps.
- Communicating with stakeholders. A single curve is far easier to discuss than a list of coefficients, and ICE lines show whether the story is uniform or varies by person.
- Catching spurious features. Weird, jagged, or implausible relationships on ID-like or time-leaking features often show up visually before they show up in metrics.
A Practical Workflow
- Start with the top features by importance, since plotting everything creates noise.
- Plot PDP with ICE lines for each, centered if baselines vary.
- Check data density along the x-axis and the correlation structure with other features.
- Use two-way PDPs for suspected interactions.
- Compare against SHAP or ALE for features where correlation could distort the picture.
- Review the curves with a domain expert, and save the plots with model version and data sample details for later audits — the same documentation habit the model cards guide describes.
Common Pitfalls
- Reading a flat PDP as "this feature doesn't matter" without looking at ICE lines.
- Trusting curves in regions with few data points, or for strongly correlated features, without checking.
- Overinterpreting small wiggles without looking at the y-axis scale.
- Plotting hundreds of ICE lines until nothing is readable, instead of subsampling.
- Treating PDPs as causal effects.
- Computing plots on unrepresentative data, so the averages describe the wrong population.
Recommended Books
| Cover | Book | Description | Get it |
|---|---|---|---|
![]() |
Interpretable Machine Learning | the chapter treatment of PDP, ICE, and ALE this tutorial condenses. | View on Amazon |
![]() |
The Elements of Statistical Learning | where Friedman's H-statistic comes from. | View on Amazon |
![]() |
Designing Machine Learning Systems | how model inspection fits into the production loop. | View on Amazon |
Unlock AI That Actually Works
Get lifetime access to GPT-6 Astra, Claude Fable 5.1, Gemini 3.5, Grok 4.5, and more — all in one platform. Build websites, apps, videos, content, and digital products from a single command. No monthly fees. No tool-hopping.
Click here to get GPTAstra Max now — one-time payment, lifetime access.
Frequently Asked Questions
What is the difference between a PDP and an ICE plot?
A partial dependence plot shows the average prediction as you sweep one feature across its range with every other feature left as observed, while an ICE plot skips the averaging and draws one line per data point showing how that individual's prediction changes as the feature varies. The relationship is direct: the PDP is the average of the ICE lines — scikit-learn's fast recursion method computes the PDP by implicitly averaging ICE curves, which is why it can't produce individual lines. Use both: the PDP summarizes, and the ICE lines show whether that summary describes everyone.
Why can a flat PDP be misleading?
Because averaging can cancel opposite effects. If half your customers respond to a discount with rising purchase probability and the other half respond with falling probability, the average is flat even though the model is reacting strongly in opposite directions — this happens whenever a feature interacts with another feature. ICE lines expose it: if they fan out, cross, or split into groups, an interaction exists and the PDP's single curve is misleading; if they run roughly parallel, the average is a fair description. Never conclude a feature is irrelevant from a flat PDP alone.
What is the main limitation of partial dependence plots?
PDPs rest on an independence assumption: overwriting a feature with a grid value across all rows creates combinations that may never occur in real data, such as setting age to 25 for a retiree's other attributes, so the model is evaluated where it has little training data and its behavior can be arbitrary. Practical responses: check pairwise correlations first, treat regions with sparse data ticks as unreliable, consider ALE plots (Accumulated Local Effects, by Apley and Zhu), which use only local realistic neighborhoods, and cross-check against SHAP dependence plots with different assumptions — treat disagreement as a reason to investigate.
How do you read a partial dependence plot?
Check five things: shape (monotonic rises, plateaus, thresholds, U-shapes; a sudden step often marks a learned split), scale (read the y-axis before trusting a wiggle — 0.2 units on a target spread of 100 means little), data density (the tick marks along the x-axis show deciles; sparse ticks mean the curve rests on few examples), edges (the default grid covers the 5th to 95th percentiles, trimmed via percentiles or custom_values), and ICE spread (wide spread means other features change this feature's effect — look for subgroups or try a two-way plot).
Wrapping This Up
PDPs show the average effect of a feature on a model's predictions, and ICE plots show the same relationship for each individual, with the PDP being the average of the ICE lines. Together they reveal shape, thresholds, and interactions that importance scores hide. Centering ICE lines highlights differences in effect, two-way PDPs expose pairwise interactions, and density marks tell you where the curve rests on real data.
Are they the final word on how your model works? No, correlated features can make them extrapolate, and they describe the model's behavior instead of reality. But with a quick check on correlations and a cross-check against ALE or SHAP, they remain one of the fastest ways to see what a model has actually learned from a feature.


