Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Repair Windows errors before they cause bigger problemsFix Now →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
For logits shaped (batch, classes), use dim=1 when you want one probability distribution per example. During classification training, pass the raw logits—not softmax probabilities—to CrossEntropyLoss. Use log_softmax when you need log probabilities directly.
What does dim mean in PyTorch softmax?
dim selects the tensor axis across which PyTorch normalizes values. Softmax exponentiates values and divides each by the sum of exponentials along that axis, so each slice along the selected dimension is normalized independently. The result is between 0 and 1, and each such slice sums to 1. See the PyTorch softmax documentation.
Choose the dimension that indexes the mutually exclusive classes. For a two-dimensional batch where rows are examples and columns are classes, that is dimension 1:
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →probabilities = torch.softmax(logits, dim=1)
For logits shaped (N, C, H, W), dimension 1 is the class axis; applying softmax(logits, dim=1) gives a distribution over classes at each spatial location.
#1 Best Overall
What is the difference between log_softmax and softmax?
softmax returns probabilities. log_softmax returns their logarithms, which are log probabilities and therefore are usually zero or negative. If you need log probabilities, use the dedicated operation:
log_probabilities = torch.nn.functional.log_softmax(logits, dim=1)
PyTorch documents that applying softmax and then taking a logarithm is slower and numerically unstable compared with log_softmax, which uses an alternative formulation to compute the output and gradient correctly. For a negative-log-likelihood workflow, pair log_softmax with NLLLoss. See the PyTorch log_softmax documentation.
Rank #2
Should I apply softmax before CrossEntropyLoss?
No. Pass unnormalized logits directly to CrossEntropyLoss; do not apply softmax first. For class-index targets, the loss is equivalent to applying LogSoftmax followed by NLLLoss. PyTorch’s CrossEntropyLoss documentation describes the accepted input shapes and target forms.
# logits: (batch, classes); targets: class IDs, shape (batch)
loss_fn = torch.nn.CrossEntropyLoss()
loss = loss_fn(logits, targets)
# Convert to probabilities separately when needed for inference or reporting
probabilities = torch.softmax(logits, dim=1)
This example assumes dimension 1 is the class axis. For a different layout, identify the axis containing classes and choose dimensions and target shapes accordingly.
Rank #3
Which input and target shapes does CrossEntropyLoss accept?
The loss accepts an unbatched class vector, a batch-by-class matrix, or higher-dimensional inputs in which dimension 1 is the class axis.
| Logits shape | Class-index target shape | Notes |
|---|---|---|
(C) |
Scalar class ID | One unbatched example. |
(N, C) |
(N) |
One class ID per example. |
(N, C, d1, …, dK) |
(N, d1, …, dK) |
Dimension 1 is classes; targets omit that axis. |
Class-index targets
Use class IDs in the range [0, C), except that a configured ignore_index may be used for ignored targets. For spatial logits, target dimensions correspond to the non-class dimensions, such as (N, H, W) for logits shaped (N, C, H, W).
Rank #4
Class-probability targets
Probability targets have the same shape as the logits, including the class dimension, and each target should be a valid probability distribution. PyTorch does not strictly check that these constraints hold; invalid target values can produce misleading loss values and unstable gradients. Class-index targets generally allow a more optimized computation, so use probability targets when soft or blended labels are actually needed. Details are in the CrossEntropyLoss target documentation.
Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesWhich CrossEntropyLoss options affect the result?
reductioncontrols output aggregation:'none'returns individual losses,'sum'sums them, and'mean'is the documented default.weightsupplies class weights.ignore_indexapplies to class-index targets.label_smoothingenables label smoothing.
The precise meaning of 'mean' depends on target form. For class indices, the documented mean accounts for class weights and ignored targets; for probability targets, it divides summed element losses by the number of loss elements. Consult the PyTorch loss documentation for the behavior of the release you use.
Common mistakes to avoid
- Normalizing the wrong axis: choosing the batch axis for a batch-by-class tensor makes values sum across examples instead of across classes. Select the class axis.
- Applying softmax before cross-entropy: pass logits to
CrossEntropyLoss; convert them to probabilities separately only when needed. - Taking
log(softmax(x)): uselog_softmaxwhen log probabilities are needed to avoid the slower, less numerically stable separate operations. - Giving class IDs a class axis: class-index targets omit the class dimension; probability targets have the same shape as logits.
- Assuming probability targets are validated: ensure they are valid distributions yourself because PyTorch does not strictly enforce those constraints.
These APIs are documented across PyTorch’s main functional documentation and stable documentation labeled 2.14 for CrossEntropyLoss. Documentation and behavior can change, so check the documentation corresponding to the PyTorch release used by your project.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

