Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

To calculate a PyTorch nn.Conv2d output shape, keep the batch and channel dimensions separate, then apply the documented formula to height and width. For each spatial axis, the result is the floor of the padded input size minus the dilated kernel extent, divided by stride, plus one. The output channel count is simply out_channels.

What nn.Conv2d expects and returns

PyTorch’s Conv2d applies a 2D convolution over an input with multiple planes. Its operation is implemented as cross-correlation, with a learned bias added for each output channel. The documented API is PyTorch nn.Conv2d.

A batched input has shape (N, C_in, H_in, W_in) and produces (N, C_out, H_out, W_out). An unbatched input may instead have shape (C_in, H_in, W_in), producing (C_out, H_out, W_out). Here, N is batch size, C_in must equal the layer’s in_channels, and C_out is the configured out_channels.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to calculate the output height and width

For tuple-valued spatial parameters, where each pair is ordered as (height, width), calculate each axis independently:

H_out = floor((H_in + 2*padding[0] - dilation[0]*(kernel_size[0] - 1) - 1) / stride[0] + 1)
W_out = floor((W_in + 2*padding[1] - dilation[1]*(kernel_size[1] - 1) - 1) / stride[1] + 1)

For scalar values such as kernel_size=3, use the same value for both axes. The floor operation means a fractional result is rounded down; stride does not guarantee that the input dimensions will divide evenly into the output positions.

Worked example

Consider an input shaped (20, 16, 50, 100) and this documented configuration: nn.Conv2d(16, 33, (3, 5), stride=(2, 1), padding=(4, 2), dilation=(3, 1)).

  • Height: floor((50 + 2*4 - 3*(3-1) - 1) / 2 + 1) = 27.
  • Width: floor((100 + 2*2 - 1*(5-1) - 1) / 1 + 1) = 100.

The resulting shape is (20, 33, 27, 100). These dimensions follow from the documented formula and configuration.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What each Conv2d parameter controls

The module’s signature is:

nn.Conv2d(
    in_channels,
    out_channels,
    kernel_size,
    stride=1,
    padding=0,
    dilation=1,
    groups=1,
    bias=True,
    padding_mode="zeros",
    device=None,
    dtype=None,
)
Parameter What it controls Effect to keep in mind
in_channels Number of channels in the input. Must match the input tensor’s channel dimension.
out_channels Number of filters and output channels. Sets the output channel dimension and affects weight and bias counts.
kernel_size Height and width of the convolution window. A pair permits non-square kernels; a larger effective kernel can reduce spatial output unless padding compensates.
stride Distance the window moves between positions. Values larger than one generally produce fewer positions; height and width can use different strides.
padding Implicit padding on each side of each spatial axis. Numeric pairs specify height then width; string options are 'valid' and 'same'.
dilation Spacing between kernel points. Raises the effective kernel extent without increasing the kernel’s stored weight dimensions.
groups Partitions input and output channel connections. Must divide both channel counts; larger grouping reduces connections and weight count.
bias Whether to learn a bias for each output channel. When false, there are no bias parameters.
padding_mode How numeric padding is filled. Documented modes are 'zeros', 'reflect', 'replicate', and 'circular'.
device and dtype Optional device and data type for the layer’s parameters. They do not alter the output-size formula.

For kernel_size, stride, numeric padding, and dilation, an integer applies to both height and width. A two-value tuple gives separate values in height-then-width order.

Padding choices

  • padding='valid' means no padding.
  • padding='same' pads to preserve the input’s height and width, but only supports stride 1.
  • With numeric padding, each specified amount is applied on both sides of its axis. For example, height padding 4 adds four positions at the top and four at the bottom for the shape calculation.

How many learnable parameters does Conv2d have?

The weight tensor has shape (out_channels, in_channels / groups, kernel_height, kernel_width). If bias is enabled, its shape is (out_channels,). Therefore:

parameters = out_channels * (in_channels / groups) * kernel_height * kernel_width
           + (out_channels if bias else 0)

For Conv2d(16, 33, 3, stride=2), defaults include groups=1 and bias=True, so the count is 33 * 16 * 3 * 3 + 33 = 4,785. Stride changes the number of output positions, not the number of learned weights.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

How groups change channel connectivity

Both in_channels and out_channels must be divisible by groups. With groups=1, each output channel can use every input channel. With groups=2, the input and output channels are divided into two channel groups, and connections stay within those groups.

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A depthwise convolution is the special case where groups == in_channels and out_channels == K * in_channels for a positive integer K. Each input channel is processed in its own group, with K output channels per input channel.

Complete shape example in Python

This snippet uses the same layer and input dimensions as the worked calculation:

import torch
from torch import nn

layer = nn.Conv2d(
    in_channels=16,
    out_channels=33,
    kernel_size=(3, 5),
    stride=(2, 1),
    padding=(4, 2),
    dilation=(3, 1),
)
x = torch.randn(20, 16, 50, 100)
y = layer(x)
print(y.shape)  # (20, 33, 27, 100), from the documented shape formula

The expected dimensions in the comment are calculated from the API formula.

Implementation details that can matter

The Conv2d reference documents support for TensorFloat32 and complex data types. It also notes that on certain ROCm devices, float16 inputs use different precision for backward computation. These are conditional backend details, not guarantees that apply identically to every device.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The PyTorch functional conv2d reference notes that some CUDA/cuDNN circumstances may select a nondeterministic algorithm for performance. When determinism is preferred, it points to torch.backends.cudnn.deterministic = True; enabling it may reduce performance. This is an implementation option, not a claim that every convolution is nondeterministic.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.