Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

In PyTorch, a direct tensor calculation and a class derived from nn.Module can perform exactly the same model computation. The difference is how the model’s parameters and other state are organized and exposed: a module registers them so PyTorch can discover them for optimization, device changes, composition, and saving or loading state.

What is the difference between raw tensors and nn.Module?

Consider an affine model: multiply an input by a weight matrix, then add a bias. The arithmetic is y = x @ weight + bias in either implementation. A raw-tensor version keeps references to the tensors and performs the operation directly. A module version places the same operation in forward and makes learnable values registered attributes.

Direct tensor implementation

import torch

weight = torch.randn(3, 2, requires_grad=True)
bias = torch.randn(2, requires_grad=True)

x = torch.randn(4, 3)
y = x @ weight + bias
loss = y.square().mean()
loss.backward()

optimizer = torch.optim.SGD([weight, bias], lr=0.01)
optimizer.step()

Autograd can calculate gradients for these tensors because they require gradients. nn.Module is not a prerequisite for autograd. The author must, however, pass the intended tensors to the optimizer and manage any model state that should be saved, restored, or moved to another device or dtype.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

The same computation as a module

import torch
from torch import nn

class Affine(nn.Module):
    def __init__(self):
        super().__init__()
        self.weight = nn.Parameter(torch.randn(3, 2))
        self.bias = nn.Parameter(torch.randn(2))

    def forward(self, x):
        return x @ self.weight + self.bias

model = Affine()
x = torch.randn(4, 3)
y = model(x)
loss = y.square().mean()
loss.backward()

optimizer = torch.optim.SGD(model.parameters(), lr=0.01)
optimizer.step()

The module adds registration and a standard interface; it does not change the affine function. PyTorch’s API defines torch.nn.Module as the “Base class for all neural network modules.” The same documentation recommends building models by subclassing it.

#1 Best Overall
Sale
Deep Learning (Adaptive Computation and Machine Learning series)
  • Language Published: English
  • Binding: hardcover
  • It ensures you get the best usage for a longer period

How does nn.Module track parameters?

Assigning an nn.Parameter to a module attribute registers it as a learnable parameter. The module can then expose it through parameters() or named_parameters(). A plain tensor attribute is not automatically registered as a parameter merely because it is stored on a module.

This distinction matters when constructing an optimizer. With the module, torch.optim.SGD(model.parameters(), lr=0.01) obtains the registered parameters from the model. With raw tensors, the optimizer must be given those tensors explicitly, such as [weight, bias]. Forgetting one means that tensor will not be updated by that optimizer.

For common layers, built-in modules such as nn.Linear provide registered parameters and a ready-made forward computation. Custom modules are useful when the computation or composition does not fit a built-in layer.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How do child modules and buffers fit in?

Child modules

A module can contain other modules. Assigning a child module to an attribute registers it with the parent, allowing parent-level traversal of parameters and state. This is how a model assembled from layers can expose its nested components through one top-level object. Call super().__init__() in the parent’s constructor before assigning child modules.

Buffers

Not all module state is learnable. A buffer is state that should be associated with a module but is not treated as a parameter; BatchNorm running statistics are a typical example. Persistent buffers are included in the module’s state_dict, while non-persistent buffers are excluded. Both kinds are affected by module-wide device and dtype changes through to().

What do module device changes and checkpoints manage?

Calling a module operation such as model.to(...) applies to its registered parameters and buffers, including those in registered child modules. In a raw-tensor implementation, the author must keep track of the relevant tensors and move or convert them as needed.

A module’s state_dict() contains its parameters and persistent buffers, keyed by their names in the module hierarchy. It is a shallow copy whose values refer to the module’s parameters and buffers; by default, the returned tensors are detached from autograd. The dictionary is useful for saving and restoring state, but it is not the Python class or executable model architecture.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

To restore a state dictionary, construct a compatible module and load the state into it. With strict loading, the checkpoint keys must match the keys expected by the module.

Best Value
Sale
Deep Learning: A Visual Approach
  • Deep Learning: A Visual Approach
  • No Starch Press
  • ABIS BOOK
model = Affine()
torch.save(model.state_dict(), "affine_state.pt")

restored = Affine()
state = torch.load("affine_state.pt", weights_only=True)
restored.load_state_dict(state)

For raw tensors, saving and restoring equivalent state is possible, but the code must define how the tensors are collected, named, and copied back. The module supplies a conventional state traversal and loading interface.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Which approach should you use?

Concern Raw tensors nn.Module
Where weight and bias live Explicit tensor variables or references kept by your code. Registered attributes, commonly nn.Parameter objects.
Optimizer input Pass the intended tensors directly, for example [weight, bias]. Use model.parameters() to enumerate registered parameters.
Nested components Your code must organize and traverse them. Assign child modules as attributes; the parent registers them recursively.
Device and dtype changes Move or convert tracked tensors yourself. Module operations such as to() apply to registered parameters and buffers.
Saving and restoring model state Define your own state organization and restoration logic. Use state_dict() and load_state_dict() with a compatible module.

For a small experiment that only needs a few tensor operations, direct tensors can be clear and entirely valid. Use nn.Module when you want the framework’s standard parameter and state management, or when building a model from reusable or nested components. Neither representation is inherently a faster implementation: the comparison here is about organization and framework integration, not a measured performance difference.

The examples follow PyTorch 2.14 documentation. Exact APIs and serialization behavior can vary between releases; consult the documentation for the version installed in your environment.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Quick Recap

SaleBestseller No. 1
Deep Learning (Adaptive Computation and Machine Learning series)
Deep Learning (Adaptive Computation and Machine Learning series)
Language Published: English; Binding: hardcover; It ensures you get the best usage for a longer period
$51.51
SaleBestseller No. 2
SaleBestseller No. 5
Deep Learning: A Visual Approach
Deep Learning: A Visual Approach
Deep Learning: A Visual Approach; No Starch Press; ABIS BOOK
$66.76

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.