Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

torch.nn.MSELoss squares the difference between each prediction and its corresponding target. Its default, reduction='mean', averages those squared differences across every tensor element. Use 'sum' to add them or 'none' to keep the elementwise loss. For an ordinary comparison, make prediction and target the same shape: a broadcastable mismatch can calculate a loss between unintended values.

What does PyTorch MSELoss return?

For each corresponding element, mean squared error is:

(input - target) ** 2

The PyTorch MSELoss API accepts tensors with any number of dimensions and documents input and target as having the same shape. What the call returns depends on its reduction setting.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do the three reduction options differ?

Let N be the total number of elements in the loss tensor. The reduction determines whether those per-element squared differences remain separate, are added, or are averaged.

reduction Result When it is useful
'none' Elementwise squared differences, with the input and target shape When you need to inspect, mask, or further aggregate individual losses
'sum' The sum of all elementwise squared differences When the total loss, rather than an element average, is required
'mean' (default) The sum divided by N When you want the average squared error over all elements

For tensors shaped [batch, channels, height, width], the default mean includes every batch, channel, height, and width element. It is not automatically a mean of per-sample means. If your objective needs samples or features to carry different weights, define that aggregation explicitly.

How do you use nn.MSELoss and F.mse_loss?

nn.MSELoss is a module you can create once and call on predictions and targets. torch.nn.functional.mse_loss is a direct function call. Both support 'none', 'sum', and 'mean', with 'mean' as the default.

Reusable module

import torch.nn as nn

criterion = nn.MSELoss(reduction='mean')
loss = criterion(prediction, target)

Functional call

import torch.nn.functional as F

loss = F.mse_loss(prediction, target, reduction='mean')

The functional API documentation also lists an optional weight argument. Because the available documentation is for PyTorch main, check the documentation for your installed version before relying on that argument.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The API references list legacy size_average and reduce arguments as deprecated; if specified, they override reduction for now. Use reduction in new code.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

What should you do when prediction and target shapes differ?

For an ordinary element-by-element MSE comparison, reshape or otherwise align the tensors so each prediction is paired with its intended target. PyTorch’s broadcasting rules can allow some differently shaped tensors to participate in elementwise operations, but a valid broadcast does not guarantee a meaningful pairing.

Why [B, 1] and [B] can be a problem

Broadcasting aligns dimensions from the right. If one tensor has shape [B, 1] and the other has shape [B], the dimensions can expand to [B, B]. Instead of comparing each row with its matching target, the operation can compare every row against every target. If each prediction should match one target, make the target shape [B, 1] as well.

Check shapes before computing the loss

  1. Inspect prediction.shape and target.shape before calling the loss.
  2. Choose an explicit reshape or unsqueeze only when it represents the intended sample, channel, or feature layout.
  3. Check the resulting shapes, especially if you deliberately expand a dimension or use broadcasting.
  4. Do not assume tensors are interchangeable just because they contain the same number of elements. Equal element counts do not establish that their dimensions align as intended.

PyTorch’s broadcasting documentation notes that older pointwise behavior that flattened some equal-element-count inputs was deprecated. Explicit dimensions make the comparison easier to reason about and help avoid unintended broadcasted loss values.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.