Quick wins for a faster PC:
Fix the driver behind crashes, sound loss and screen glitchesFind Drivers →Clear out junk files and repair common Windows errorsFree Scan →iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
In this article, EANet means the External Attention Transformer from the official Keras image-classification example. That example trains a patch-based transformer on CIFAR-100 and replaces standard self-attention with external attention. The model is small enough to study and run, and it shows how the pieces fit together in Keras code. The acronym is also used for unrelated architectures in other papers, so when you search, keep the Keras example in mind as your reference. The source is the Keras EANet example page, written by ZhiYong Chang.
Which EANet this article covers
The Keras page titles its example “Image classification with EANet (External Attention Transformer).” Everything below refers to that model and its configuration. If you came across EANet in a paper that uses a different design, the attention mechanism and the Keras code described here will not necessarily match it.
The dataset and what the model predicts
The example uses CIFAR-100. It has 50,000 training images and 10,000 test images, each 32×32 pixels with three RGB channels, spread across 100 classes. The model outputs one probability per class through a 100-way softmax layer, so the labels are one-hot encoded to 100 classes before training.
Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Scan for outdated or missing drivers - takes under a minute3Repair Windows errors before they cause bigger problemsWhat external attention changes
A standard transformer block computes self-attention, where every patch attends to every other patch in the same image. The Keras page describes EANet’s alternative this way:
#1 Best Overall
“EANet introduces a novel attention mechanism named external attention, based on two external, small, learnable, and shared memories, which can be implemented easily by simply using two cascaded linear layers and two normalization layers.”
In practical terms, the patches do not attend to each other directly. Instead, each image’s patch features are compared against a small set of learned memory vectors that are shared across all images. Those memories are learned during training, so the model learns what to look for rather than computing pairwise relations from scratch. Because the memories are shared, the same learned parameters are used for every input.
Rank #2
The model pipeline, step by step
- Data augmentation. Training images are augmented before they reach the network.
- Patch extraction. Each 32×32 image is cut into 2×2 patches. That gives 16 × 16 = 256 patches per image.
- Patch embedding. Each patch is projected into a 64-dimensional embedding.
- Transformer encoder blocks. Eight blocks are stacked. Each uses external attention with four heads, in place of self-attention.
- Global average pooling. The sequence of patch representations is averaged into a single vector per image.
- Classification. A dense layer with 100 outputs and a softmax produces the class probabilities.
Because the patch size must divide the image size evenly, 2×2 patches on a 32-pixel image produce exactly 256 patches. Changing the input resolution changes this count.
Free tools Windows power users keep installed
One-click scans. No signup required.
Configuration used in the example
The values below are the ones the example sets. They are a working configuration for this demonstration, not general recommendations for every dataset.
| Setting | Value in the example | What it controls |
|---|---|---|
| Patch size | 2×2 | Size of each image patch |
| Patches per image | 256 | Sequence length fed to the encoder |
| Embedding dimension | 64 | Width of each patch representation |
| Attention heads | 4 | Number of attention heads per block |
| Transformer blocks | 8 | Depth of the encoder |
| Attention and projection dropout | 0.2 | Regularization inside each block |
| Batch size | 128 | Images per gradient update |
| Epochs | 50 | Full passes over the training set |
| Learning rate | 0.001 | Optimizer step size |
| Weight decay | 0.0001 | Penalty on weight magnitude |
| Label smoothing | 0.1 | Softens the one-hot targets in the cross-entropy loss |
The example also holds out a validation split from the training data. The page describes this setting but the value is set in the code, so check the script before you copy it.
Cost: self-attention versus external attention
The Keras page gives an asymptotic account of the cost of each attention type. Here, d is the feature dimension, N is the number of patches, and S is the size of the external memory, which is a hyperparameter you choose.
| Attention type | Complexity stated on the page | Dependence on sequence length |
|---|---|---|
| Self-attention | O(d · N²) | Quadratic in N |
| External attention | O(d · S · N) | Linear in N for fixed S |
This is a theoretical scaling statement, not a measured runtime. The page does not report wall-clock timings or accuracy numbers for the two attention types, so do not read this table as a promise of speed or accuracy. The gain matters most when N is large; with only 256 patches, the practical difference on your hardware may be small.
Implementation outline
The example imports keras, layers, and ops, loads CIFAR-100, one-hot encodes the labels, and sets the input shape to 32, 32, 3. The simplified fragment below shows those parts; copy the full script from the official page for the complete model and training loop.
Best Value
import keras
from keras import layers
from keras import ops
(x_train, y_train), (x_test, y_test) = keras.datasets.cifar100.load_data()
input_shape = (32, 32, 3)
num_classes = 100
The rest of the script defines the augmentation layers, the patch extraction and embedding layers, the external attention block, and the pooling and softmax head. Each block is built from two linear layers and two normalization layers around the shared memories, as described above.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Training setup
The example uses categorical cross-entropy with label smoothing, a weight decay term, a fixed learning rate, a validation split, a batch size of 128, and 50 epochs. Those values produce the example’s configuration. They are not tuned results, and the page does not report a final accuracy that you can compare your run against.
Running the example on your own setup
The page was created on 19 October 2021 and last modified on 18 July 2023. It does not pin a Keras version, so the code may need small changes on a current release. Before running it:
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.
- Install Keras and confirm which backend you are using.
- Expect the first run to download CIFAR-100 automatically.
- Run the full 50-epoch schedule on a GPU if you can; training on a CPU will take much longer.
- If you switch datasets, update the input shape and class count together, and make sure the patch size still divides the image width and height.
- If you change the image size, recalculate the patch count (for example, a 64×64 image with 2×2 patches gives 1,024 patches) and expect the cost to rise.
Keep the learning rate, weight decay, and label smoothing as starting points, and adjust them only after you have a baseline run on your own data.
Quick Recap
The Bottom Line
“”
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

