What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
To visualize what a CNN responds to, choose between two views: a feature map shows where a channel responds to a particular image, while a filter visualization creates a synthetic image that strongly excites a chosen channel. Use feature maps to inspect a real input; use activation maximization to probe what a learned channel prefers. Neither view, on its own, tells you which image regions caused a class prediction.
What is the difference between a filter and a feature map?
A convolutional layer has learned filters (also called kernels) that produce output channels. In common CNN usage, “visualize a filter” can mean displaying the kernel’s weights, but that is not the same as seeing what input pattern activates the channel. For interpreting a channel’s learned response, activation maximization synthesizes an input that raises its activation.
- Feature map: the spatial pattern of one output channel for a supplied image. Bright areas indicate locations where that channel responds strongly.
- Activation-maximization image: a synthetic probe optimized to excite a selected channel. It is not a photograph retrieved from the training data.
- Class-attribution map: a visualization of input regions relevant to a particular prediction, using methods such as Grad-CAM or saliency. This answers a different question from either channel visualization.
A channel can respond at several locations in one feature map. A synthesized input, by contrast, does not show where that channel fired in a particular photograph.
How do you inspect feature maps for a real image?
1. Find the convolutional layer and its output channels
Inspect the model summary or layer list and choose a convolutional layer. Record its exact name and output-channel count. The channel index used in code is zero-based, so a layer with 64 channels has indices 0 through 63. The example below assumes a channels-last model and one input image in NHWC layout: batch, height, width, channels.
The Tool Desk
Outbyte Driver Updater FREEScan for outdated or missing drivers - takes under a minuteDriver Scan →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →#1 Best Overall
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
2. Make an extractor for that layer
In Keras, create a model that shares the original model’s input and returns the chosen layer’s output. This lets you obtain intermediate activations without changing the trained network:
layer = model.get_layer(name=layer_name)
feature_extractor = keras.Model(
inputs=model.inputs,
outputs=layer.output,
)
Use the same model and weights that produced the prediction you want to inspect. If the selected layer has multiple outputs or is not a convolutional layer, adapt the output selection accordingly.
Rank #2
3. Preprocess the image as the model expects
Apply the exact input preprocessing used during training or inference, including resizing, channel order, value range, and any model-specific normalization. Then pass the resulting single-image batch to the extractor. A preprocessing mismatch can change the activations, making the visualization misleading.
4. Plot selected channels with a shared scale
The extractor’s output for a channels-last convolutional layer is typically a spatial tensor with one plane per channel. Select the planes you want, then plot them as grayscale maps or a tiled grid. This Matplotlib example plots the first 16 channels and uses one shared color range, so a brighter pixel has a comparable displayed value across those panels:
Free tools Windows power users keep installed
One-click scans. No signup required.
import matplotlib.pyplot as plt
# image_batch is one preprocessed image with shape (1, height, width, channels).
features = feature_extractor(image_batch, training=False).numpy()
feature_maps = features[0] # height, width, channels
channel_indices = list(range(min(16, feature_maps.shape[-1])))
vmin = min(feature_maps[:, :, i].min() for i in channel_indices)
vmax = max(feature_maps[:, :, i].max() for i in channel_indices)
fig, axes = plt.subplots(4, 4, figsize=(10, 10))
for ax, channel_index in zip(axes.flat, channel_indices):
ax.imshow(
feature_maps[:, :, channel_index],
cmap="gray",
vmin=vmin,
vmax=vmax,
)
ax.set_title(f"Channel {channel_index}")
ax.axis("off")
for ax in axes.flat[len(channel_indices):]:
ax.axis("off")
plt.tight_layout()
plt.show()
For a single channel, plot just its two-dimensional plane. If you normalize every panel separately, each map may look high-contrast even when its absolute responses are much smaller than the others. Use a shared scale when comparing response magnitudes; per-map scaling is useful only when the goal is to inspect each map’s spatial pattern independently. Keep the color bar or document the scale if the figure will be compared or reused.
How do you visualize what a learned channel prefers?
Use gradient ascent on the input to increase the mean activation of a selected channel. The optimization objective should target the channel’s activation rather than a class score. In the pattern below, the crop excludes two pixels at each edge of the activation map, reducing edge artifacts in the objective:
Rank #4
layer = model.get_layer(name=layer_name)
feature_extractor = keras.Model(
inputs=model.inputs,
outputs=layer.output,
)
# img is a trainable, batched image in the model's expected input space.
# Initialize it from a neutral or random image using the model's input shape.
for _ in range(steps):
with tf.GradientTape() as tape:
activation = feature_extractor(img, training=False)
filter_activation = activation[:, 2:-2, 2:-2, filter_index]
loss = tf.reduce_mean(filter_activation)
grads = tape.gradient(loss, img)
grads = tf.math.l2_normalize(grads)
img.assign_add(learning_rate * grads)
# Convert the optimized image back to displayable RGB values using the
# inverse of the model-specific preprocessing, then clip for display.
Here, layer_name, filter_index, steps, learning_rate, and img must match your model and input setup. The image variable must be watched by the gradient tape (a tf.Variable is watched automatically). The input’s value range and color preprocessing are model-specific: optimize in the representation the model expects, and apply the corresponding inverse preprocessing before displaying RGB. Clipping directly to 0–255 or 0–1 is appropriate only if that is the display-space range after the inverse transform.
The result depends on the starting image, objective, preprocessing, optimization settings, and any regularization you add. Different choices can produce different patterns for the same channel. Treat the result as a probe of the channel’s preferences, not as a unique or literal depiction of what the CNN “sees.”
Best Value
How should you compare layers and interpret the images?
Inspect channels from early, middle, and late convolutional layers rather than drawing conclusions from a single layer. Early layers often make edge-, color-, or texture-like responses easier to recognize; deeper layers often combine lower-level signals into more complex patterns. This is a useful tendency, not a rule that every channel follows. Keras describes the progression as a “modular-hierarchical decomposition of its visual space.”
- For a feature map, interpret location as well as intensity: high responses at multiple positions mean the channel responded in multiple parts of that input.
- For an optimized image, remember that the visualization is synthetic and shaped by the optimization choices. It is not evidence that the training set contained that exact pattern.
- When comparing feature-map magnitudes across channels, keep the display scale consistent. Independent normalization can conceal magnitude differences.
- Keep the layer name, channel index, source image, preprocessing, weights, iteration count, and random seed with the figure so the visualization can be reproduced.
Which visualization should you use for your question?
| Method | Question it addresses | Requires a real input image? | Spatially localizes the response? | Main consideration |
|---|---|---|---|---|
| Feature-map grid | Where does a channel respond in this image? | Yes | Yes, within the selected channel’s map | Preprocessing and color scaling affect interpretation. |
| Activation maximization | What synthetic input pattern raises this channel’s activation? | No; it optimizes a synthetic input | No, not for a supplied photograph | Results depend on initialization, objective, preprocessing, optimization, and regularization. |
| Grad-CAM, Grad-CAM++, Score-CAM, or LayerCAM | Which input regions are relevant to a selected class prediction? | Yes | Yes, as an input-region attribution map | These methods target a prediction, not a general description of a filter. |
| Saliency map | How does the prediction relate to input-level sensitivity? | Yes | Yes, at the input level | It answers an attribution question, not which pattern a channel prefers. |
Grad-CAM, Grad-CAM++, Score-CAM, Faster-Score-CAM, LayerCAM, vanilla saliency, and SmoothGrad are available in the tf-keras-vis toolkit. Choose among them based on whether you need a class-specific region explanation, rather than treating them as substitutes for a feature-map grid or activation maximization.
What should you check if a visualization looks wrong?
- Blank or nearly uniform maps: verify that the image preprocessing and channel order match the model’s expected input, and confirm that the selected layer and channel index are correct.
- Unexpected colors or scale: check the input value range and display range. For activation maps, compare channels using a shared scale rather than normalizing each independently.
- Optimization produces edge patterns: use the cropped activation objective shown above to reduce the influence of border responses; initialization and optimization settings can also affect the result.
- A filter image does not resemble a recognizable object: that does not by itself mean the channel is useless. The optimized pattern is a synthetic probe, and a channel may respond to lower-level or distributed features rather than a whole object.
- You need to explain one prediction: use a class-attribution method such as Grad-CAM or a saliency map alongside channel inspection. A filter grid alone does not identify which regions supported that class score.
Further reading
The Keras guide on visualizing what ConvNets learn provides an end-to-end gradient-ascent example using a pretrained ResNet50V2 and the intermediate layer conv3_block4_out, including a grid of 64 optimized filter images. For the research foundations of intermediate-feature visualization, see Matthew D. Zeiler and Rob Fergus, “Visualizing and Understanding Convolutional Networks” (2013 preprint; ECCV 2014 paper). Keras also points readers to Chapter 10, “Interpreting what ConvNets learn,” in Deep Learning with Python.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitches

