Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesiTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
In PyTorch, a direct tensor calculation and a class derived from nn.Module can perform exactly the same model computation. The difference is how the model’s parameters and other state are organized and exposed: a module registers them so PyTorch can discover them for optimization, device changes, composition, and saving or loading state.
What is the difference between raw tensors and nn.Module?
Consider an affine model: multiply an input by a weight matrix, then add a bias. The arithmetic is y = x @ weight + bias in either implementation. A raw-tensor version keeps references to the tensors and performs the operation directly. A module version places the same operation in forward and makes learnable values registered attributes.
Direct tensor implementation
import torch
weight = torch.randn(3, 2, requires_grad=True)
bias = torch.randn(2, requires_grad=True)
x = torch.randn(4, 3)
y = x @ weight + bias
loss = y.square().mean()
loss.backward()
optimizer = torch.optim.SGD([weight, bias], lr=0.01)
optimizer.step()
Autograd can calculate gradients for these tensors because they require gradients. nn.Module is not a prerequisite for autograd. The author must, however, pass the intended tensors to the optimizer and manage any model state that should be saved, restored, or moved to another device or dtype.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREEClear out junk files and repair common Windows errorsFree Scan →The same computation as a module
import torch
from torch import nn
class Affine(nn.Module):
def __init__(self):
super().__init__()
self.weight = nn.Parameter(torch.randn(3, 2))
self.bias = nn.Parameter(torch.randn(2))
def forward(self, x):
return x @ self.weight + self.bias
model = Affine()
x = torch.randn(4, 3)
y = model(x)
loss = y.square().mean()
loss.backward()
optimizer = torch.optim.SGD(model.parameters(), lr=0.01)
optimizer.step()
The module adds registration and a standard interface; it does not change the affine function. PyTorch’s API defines torch.nn.Module as the “Base class for all neural network modules.” The same documentation recommends building models by subclassing it.
#1 Best Overall
- Language Published: English
- Binding: hardcover
- It ensures you get the best usage for a longer period
How does nn.Module track parameters?
Assigning an nn.Parameter to a module attribute registers it as a learnable parameter. The module can then expose it through parameters() or named_parameters(). A plain tensor attribute is not automatically registered as a parameter merely because it is stored on a module.
This distinction matters when constructing an optimizer. With the module, torch.optim.SGD(model.parameters(), lr=0.01) obtains the registered parameters from the model. With raw tensors, the optimizer must be given those tensors explicitly, such as [weight, bias]. Forgetting one means that tensor will not be updated by that optimizer.
Rank #2
For common layers, built-in modules such as nn.Linear provide registered parameters and a ready-made forward computation. Custom modules are useful when the computation or composition does not fit a built-in layer.
Free tools Windows power users keep installed
One-click scans. No signup required.
How do child modules and buffers fit in?
Child modules
A module can contain other modules. Assigning a child module to an attribute registers it with the parent, allowing parent-level traversal of parameters and state. This is how a model assembled from layers can expose its nested components through one top-level object. Call super().__init__() in the parent’s constructor before assigning child modules.
Rank #3
Buffers
Not all module state is learnable. A buffer is state that should be associated with a module but is not treated as a parameter; BatchNorm running statistics are a typical example. Persistent buffers are included in the module’s state_dict, while non-persistent buffers are excluded. Both kinds are affected by module-wide device and dtype changes through to().
What do module device changes and checkpoints manage?
Calling a module operation such as model.to(...) applies to its registered parameters and buffers, including those in registered child modules. In a raw-tensor implementation, the author must keep track of the relevant tensors and move or convert them as needed.
Rank #4
A module’s state_dict() contains its parameters and persistent buffers, keyed by their names in the module hierarchy. It is a shallow copy whose values refer to the module’s parameters and buffers; by default, the returned tensors are detached from autograd. The dictionary is useful for saving and restoring state, but it is not the Python class or executable model architecture.
PC Slower Than It Used to Be?
A free scan shows the junk files, broken settings and background clutter dragging Windows down - then fixes them in one click.Free scan · Windows 10 & 11Crashes, No Sound, or Screen Glitches?
Random freezes, missing sound and display glitches usually trace back to one bad driver. Find and replace yours safely.Free scan · under a minuteTo restore a state dictionary, construct a compatible module and load the state into it. With strict loading, the checkpoint keys must match the keys expected by the module.
Best Value
model = Affine()
torch.save(model.state_dict(), "affine_state.pt")
restored = Affine()
state = torch.load("affine_state.pt", weights_only=True)
restored.load_state_dict(state)
For raw tensors, saving and restoring equivalent state is possible, but the code must define how the tensors are collected, named, and copied back. The module supplies a conventional state traversal and loading interface.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Which approach should you use?
| Concern | Raw tensors | nn.Module |
|---|---|---|
| Where weight and bias live | Explicit tensor variables or references kept by your code. | Registered attributes, commonly nn.Parameter objects. |
| Optimizer input | Pass the intended tensors directly, for example [weight, bias]. |
Use model.parameters() to enumerate registered parameters. |
| Nested components | Your code must organize and traverse them. | Assign child modules as attributes; the parent registers them recursively. |
| Device and dtype changes | Move or convert tracked tensors yourself. | Module operations such as to() apply to registered parameters and buffers. |
| Saving and restoring model state | Define your own state organization and restoration logic. | Use state_dict() and load_state_dict() with a compatible module. |
For a small experiment that only needs a few tensor operations, direct tensors can be clear and entirely valid. Use nn.Module when you want the framework’s standard parameter and state management, or when building a model from reusable or nested components. Neither representation is inherently a faster implementation: the comparison here is about organization and framework integration, not a measured performance difference.
The examples follow PyTorch 2.14 documentation. Exact APIs and serialization behavior can vary between releases; consult the documentation for the version installed in your environment.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

