In PyTorch, a Dataset defines how samples and labels are retrieved or produced; a DataLoader turns that dataset into an iterable that can batch and deliver data to a training loop. Choose a map-style dataset for indexed or keyed examples, and an iterable-style dataset for streams or sources where random access is impractical.
How PyTorch datasets and data loaders fit together
Keep data access separate from model-training code: the dataset describes where examples come from, while the loader handles iteration and, where applicable, batching and sampling. This separation makes each part easier to inspect and change. PyTorch’s beginner data tutorial demonstrates the pattern with a dataset passed to a DataLoader, then batches read from the loader in the training loop.
Dataset: retrieves or produces a sample, commonly with its label.DataLoader: supplies an iteration interface and can combine samples into batches.
Built-in datasets from PyTorch domain libraries are useful for prototyping and benchmarking. For your own data, implement a custom dataset that matches how the source can be accessed.
Choose a dataset type to match the source
| Design | How samples are obtained | Best fit | Ordering and length considerations |
|---|---|---|---|
| Map-style | By key or index through __getitem__(); __len__() may also be implemented. |
Sources with efficient retrieval of a particular example, such as indexed image and label files on disk. | Often supports sampler-based selection and loader options that expect a dataset length. If keys are not the default integer indices, use a custom sampler. |
| Iterable-style | Samples are produced by __iter__(). |
Streams, remote sources, databases, or sources where random reads are costly or impractical. | The iterable controls its own order; index-based samplers do not apply. A stable length or random-access key need not be available. |
These behaviors are described in the PyTorch data-loading documentation. The practical distinction is whether your source is naturally addressable by index or whether it should be consumed as a sequence of records.
Outdated Drivers Are Slowing You Down
One free scan finds every outdated or missing driver and matches the right update for your exact hardware.Free scan · exact hardware matchWindows Errors? Fix Them Before They Spread
Repair common Windows errors and clear accumulated junk for a smoother, more stable PC - no reinstall needed.Free scan · no reinstall#1 Best Overall
Use map-style data for indexed examples
Implement __getitem__(index) to retrieve the sample associated with an index or key. Implementing __len__() is useful because many samplers and default loader behaviors rely on knowing the dataset size. With this design, a sampler can control which indices are selected and in what order.
Use iterable-style data for streams
Implement __iter__() to yield records from the source. This is a natural fit when records arrive over time or when random reads are expensive. Since the iterable determines what comes next, do not try to configure it with an index-based sampler.
Rank #2
Pass the dataset to a DataLoader and read batches
For a map-style dataset, the loader can choose examples using shuffle or a sampler, and combine samples into batches with batch_size and collate_fn. The collate function determines how individual samples are assembled into the batch format your model expects. If the dataset size is not evenly divisible by the batch size, the last batch is smaller unless drop_last=True.
A minimal loading pattern looks like this:
from torch.utils.data import DataLoader, Dataset
class ExampleDataset(Dataset):
def __len__(self):
return len(self.records)
def __getitem__(self, index):
return self.records[index]
dataset = ExampleDataset()
loader = DataLoader(dataset, batch_size=32, shuffle=True)
for batch in loader:
# Use batch in the training step
pass
This is a structural example: define how your dataset initializes and stores or accesses records for your actual source. For iterable-style data, provide an IterableDataset and implement its iterator; ordering is then handled by that iterator rather than by shuffle or a sampler.
Free tools Windows power users keep installed
One-click scans. No signup required.
Rank #3
Use multiple workers safely, especially with iterable datasets
Set num_workers=0 to fetch data in the main process. A positive value lets the loader use worker subprocesses. For an iterable dataset, each worker gets a replica of the dataset object. If every replica reads the same source in the same way, workers can emit duplicate records instead of dividing the stream.
Shard the source so that each worker handles a distinct portion. PyTorch supports checking the active worker with get_worker_info() inside the iterable, or configuring replicas with worker_init_fn. Consult the data-loading documentation for the worker APIs and examples.
Rank #4
Tune loading performance against the real workload
There is no universally best worker count. Subprocess workers may help when storage reads or transforms are slow, but their startup, communication, and memory costs can outweigh the benefit for in-memory data or cheap operations. More workers also consume memory and may exhaust /dev/shm. Benchmark with the actual dataset, transforms, and hardware rather than treating a tutorial setting or timing as a general guarantee. The PyTorch performance tuning guide reports measurements from its own setup, not a universal result.
prefetch_factorcontrols how many batches each worker queues in advance. More prefetching can change memory use as well as delivery behavior.persistent_workers=Truekeeps workers alive between epochs instead of shutting them down and starting them again. It can help when startup or dataset initialization is expensive.- Evaluate these settings in combination with storage behavior, CPU and memory use, and the need for predictable ordering.
Consider pinned memory only when transfer is a bottleneck
pin_memory=True asks the loader to return tensors in page-locked host memory, which can improve transfers to CUDA-enabled devices in suitable workloads. PyTorch’s optimization example pairs it with .to(device, non_blocking=True). Pinning is optional, not a prerequisite for loading data, and its benefit depends on the workload; the guide’s benchmark results apply to that example’s setup.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

