LeNet-5 is a convolutional neural network for handwritten-character recognition, described by Yann LeCun, Léon Bottou, Yoshua Bengio, and Patrick Haffner in their 1998 paper, Gradient-Based Learning Applied to Document Recognition. Its original design processes a 32×32 image through alternating convolution and trainable subsampling layers, then classifies it with a radial-basis-function output layer. That last detail—and the network’s partly connected C3 layer—distinguishes the paper’s model from many simplified tutorials.
What LeNet-5 was designed to do
LeNet-5 was developed to recognize handwritten characters, especially digits. Rather than first relying on a separate, hand-designed feature-extraction pipeline, the network learns visual features from image data. Its local receptive fields, shared weights, and subsampling exploit the fact that an image has spatial structure: nearby pixels form patterns, and useful patterns may appear in different locations.
The model is part of a broader document-recognition paper, not the paper’s only subject. The authors also reviewed recognition approaches, compared methods on handwritten-digit data, and discussed graph transformer networks for training multi-module document-recognition systems as a whole. The IEEE abstract says: “Convolutional neural networks, which are specifically designed to deal with the variability of 2D shapes, are shown to outperform all other techniques.” IEEE paper record and abstract.
LeNet-5’s original layer sequence
The paper describes seven trainable layers after a 32×32 input: C1, S2, C3, S4, C5, F6, and the output layer. The labels distinguish convolutional (C), subsampling (S), and fully connected stages. The sequence is not simply a stack of modern convolution-plus-max-pooling blocks: S2 and S4 use trainable subsampling, C3 is only partially connected to S2, and the output units are RBF units rather than a softmax layer.
#1 Best Overall
| Stage | Output shape or units | Operation and connectivity |
|---|---|---|
| Input | 32×32 | Image presented to the network. |
| C1 | 6 feature maps, each 28×28 | Convolution with a 5×5 local receptive field. |
| S2 | 6 maps, each 14×14 | Trainable 2×2 subsampling. |
| C3 | 16 feature maps | Convolution with a deliberately partial set of connections to S2 maps. |
| S4 | 16 maps, each 5×5 | Subsampling that reduces spatial resolution. |
| C5 | 120 units | Each unit receives input from all S4 maps. |
| F6 | 84 units | Fully connected representation before classification. |
| Output | One RBF unit per class | Euclidean radial-basis-function outputs. |
These layer details and the original output-head description come from the authors’ paper. Full paper PDF.
Why C3 is only partly connected
C3 does not connect every output map to every S2 map. The authors selected a partial connection pattern to limit the number of connections and encourage different maps to learn complementary features. This is a deliberate architectural choice, not a missing connection in a simplified diagram.
Rank #2
Why the output is not softmax
In the original architecture, each class is represented by an RBF output unit, and the units are based on Euclidean distance. Many later educational implementations instead use a softmax classifier. A tutorial using softmax can illustrate the broad convolutional-network idea, but it is not an exact reproduction of the original output head.
How the design choices help image recognition
Local receptive fields
A convolutional unit sees a local region rather than the entire image at once. This lets early layers respond to local visual patterns such as strokes and edges, while later stages combine those responses into more complex representations.
Do these 3 things before closing this tab:
1Fix the driver behind crashes, sound loss and screen glitches2Repair Windows errors before they cause bigger problems3Scan for outdated or missing drivers - takes under a minuteRank #3
Shared weights
Within a feature map, the same detector is applied at multiple image positions. This reduces the number of independent parameters compared with learning a separate detector for every location, and lets a learned feature respond wherever it appears. It does not make the network perfectly invariant to position.
Subsampling
S2 and S4 reduce the spatial dimensions of feature maps using trainable subsampling. Lower resolution can reduce sensitivity to small shifts and make later computation more compact, but it is not a guarantee that every translation or shape variation will be handled identically.
Rank #4
What accuracy did the original LeNet-5 report on MNIST?
The paper’s modified NIST handwritten-digit experiment used 60,000 training examples and 10,000 test examples. It reported two test-error figures under different training conditions:
| Training setup in the 1998 paper | Reported test error |
|---|---|
| Regular modified-MNIST experiment, without distortion augmentation | 0.95% |
| 60,000 original patterns plus 540,000 randomly distorted instances | 0.8% |
The synthetic distortions combined translations, scaling, squeezing, and horizontal shearing. These are historical results reported by LeCun, Bottou, Bengio, and Haffner for the paper’s setup; they are not a guarantee for every implementation or a claim about a modern reproduction. The paper describes size-normalized, centered images and a 32×32 network input. The paper’s architecture and experimental results.
Best Value
How to tell an original LeNet-5 from a tutorial variant
Implementation names are not enough to establish that a model matches the 1998 architecture. When evaluating a code example or a reported error rate, check these details:
- Input and preprocessing: Does it use the paper’s 32×32 input and comparable image preparation?
- Layer widths and connections: Does it include six C1 maps, a partly connected 16-map C3, 120 C5 units, and 84 F6 units?
- Subsampling: Does it implement the paper’s trainable subsampling stages, rather than silently substituting a different pooling operation?
- Output and loss: Does it use the original Euclidean RBF class outputs or a later softmax alternative?
- Training protocol: Are the data split and any synthetic distortions specified?
- Benchmark provenance: Is the number a result reported in the 1998 paper or a separately measured modern reproduction?
Error rates should only be compared when the data split, preprocessing, augmentation, and evaluation procedure align. A lower headline number alone does not show that one architecture is better.
Why LeNet-5 remains useful to study
LeNet-5 is historically important because it provides a concrete early example of a CNN designed around image structure. Its alternating feature-extraction and subsampling stages, shared weights, selective connectivity, and distinctive classifier make it a useful case study in how architectural choices encode assumptions about the task. The original article appeared in Proceedings of the IEEE, volume 86, issue 11, in 1998. LeCun’s publication record.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Free tools Windows power users keep installed
One-click scans. No signup required.

