Use PyTorch’s map-style Dataset when a sample can be fetched by an index or key; use IterableDataset when samples are best produced as a stream. In either case, pass the dataset to DataLoader to handle batching and, where appropriate, sampling, multiprocessing, and memory pinning.
Choose the dataset style that matches your data
| Decision | Map-style Dataset |
IterableDataset |
|---|---|---|
| How samples are accessed | Look up a sample by index or key with __getitem__. |
Produce samples by yielding them from __iter__. |
| Good fit | An indexable collection that supports random access. | A stream, a source where random reads are costly, or data produced dynamically. |
| Length | __len__ is useful when the collection has a known size, but it is not required by the abstract API. |
Length may be unknown or the data may not have a natural finite end. |
| Sampling | A sampler or batch sampler can control the order of keys. | sampler and batch_sampler are incompatible. |
| Multiple workers | The main process generates indices and assigns fetches to workers. | Each worker receives a replica; configure replicas to read distinct shards to avoid duplicate samples. |
These distinctions follow PyTorch’s definitions of map-style and iterable-style datasets.
Build a map-style Dataset for indexable data
For an ordinary collection such as labeled images or rows in a table, subclass torch.utils.data.Dataset. Put setup and metadata in __init__, retrieve one sample in __getitem__, and implement __len__ when the dataset has a known size and downstream code needs it.
PyTorch’s beginner dataset tutorial demonstrates this pattern by storing annotation labels and an image directory, then retrieving a sample using its index:
Quick wins for a faster PC:
Scan for outdated or missing drivers - takes under a minuteDriver Scan →Clear out junk files and repair common Windows errorsFree Scan →Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →#1 Best Overall
from torch.utils.data import Dataset
class MyDataset(Dataset):
def __init__(self, records):
self.records = records
def __len__(self):
return len(self.records)
def __getitem__(self, index):
record = self.records[index]
feature = record["feature"]
label = record["label"]
return feature, label
The example assumes records is already indexable and each record has compatible feature and label values. In a real dataset, __getitem__ can load or transform the requested sample; keep its returned structure consistent so the loader can assemble batches.
Use IterableDataset for streams and sequential sources
Subclass torch.utils.data.IterableDataset when the source is naturally consumed in sequence, random reads are expensive, or samples are generated dynamically. Implement __iter__ as the sample-producing interface:
Rank #2
from torch.utils.data import IterableDataset
class MyStreamDataset(IterableDataset):
def __iter__(self):
for sample in read_source_in_order():
yield sample
read_source_in_order() stands for the source-specific reading logic; it is not a PyTorch function. An iterable dataset may not have a meaningful length, so do not assume that length-based operations apply as they do to an indexable collection.
Wrap either dataset in DataLoader
PyTorch describes torch.utils.data.DataLoader as the heart of its data-loading utility. A loader wraps a dataset and can provide batching, sampling options, multiprocessing, and memory pinning. For example:
Rank #3
from torch.utils.data import DataLoader
loader = DataLoader(dataset, batch_size=32, shuffle=True)
for features, labels in loader:
# use this batch in the training loop
pass
This configuration is suitable for a map-style dataset. DataLoader’s default collation can assemble batches when samples have compatible structures, such as feature-and-label tuples or dictionaries. If samples need special assembly—for example, padding variable-length sequences—provide a custom collate_fn. See the DataLoader API documentation for its options.
Configure sampling according to dataset style
For map-style datasets, DataLoader can use sequential or shuffled sampling from its configuration, or receive a custom sampler. A map-style dataset whose keys are not integral needs a custom sampler. Iterable-style datasets cannot use sampler or batch_sampler; their iteration logic determines how samples are produced. PyTorch documents these sampler requirements and restrictions.
Rank #4
Prevent duplicate samples when using multiple workers
With a map-style dataset, the main process generates indices and sends fetch work to workers. An IterableDataset, by contrast, is copied to each worker process. If every copy iterates over the same source in the same way, workers can emit duplicate data rather than divide the work.
Use worker-specific information in __iter__ to assign each replica a distinct shard—for example, a separate file, range, or partition. The shard assignment must cover the intended data without overlap. PyTorch’s multiprocessing guidance for iterable datasets explains this replica behavior.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
Check the current API for version-sensitive details
The PyTorch stable API and beginner tutorial pages reported updates on May 7, 2026. Their current documentation is the appropriate reference for option details that may change between PyTorch releases.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

