Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

In this article, EANet means the External Attention Transformer from the official Keras image-classification example. That example trains a patch-based transformer on CIFAR-100 and replaces standard self-attention with external attention. The model is small enough to study and run, and it shows how the pieces fit together in Keras code. The acronym is also used for unrelated architectures in other papers, so when you search, keep the Keras example in mind as your reference. The source is the Keras EANet example page, written by ZhiYong Chang.

Which EANet this article covers

The Keras page titles its example “Image classification with EANet (External Attention Transformer).” Everything below refers to that model and its configuration. If you came across EANet in a paper that uses a different design, the attention mechanism and the Keras code described here will not necessarily match it.

The dataset and what the model predicts

The example uses CIFAR-100. It has 50,000 training images and 10,000 test images, each 32×32 pixels with three RGB channels, spread across 100 classes. The model outputs one probability per class through a 100-way softmax layer, so the labels are one-hot encoded to 100 classes before training.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

What external attention changes

A standard transformer block computes self-attention, where every patch attends to every other patch in the same image. The Keras page describes EANet’s alternative this way:

“EANet introduces a novel attention mechanism named external attention, based on two external, small, learnable, and shared memories, which can be implemented easily by simply using two cascaded linear layers and two normalization layers.”

In practical terms, the patches do not attend to each other directly. Instead, each image’s patch features are compared against a small set of learned memory vectors that are shared across all images. Those memories are learned during training, so the model learns what to look for rather than computing pairwise relations from scratch. Because the memories are shared, the same learned parameters are used for every input.

The model pipeline, step by step

  1. Data augmentation. Training images are augmented before they reach the network.
  2. Patch extraction. Each 32×32 image is cut into 2×2 patches. That gives 16 × 16 = 256 patches per image.
  3. Patch embedding. Each patch is projected into a 64-dimensional embedding.
  4. Transformer encoder blocks. Eight blocks are stacked. Each uses external attention with four heads, in place of self-attention.
  5. Global average pooling. The sequence of patch representations is averaged into a single vector per image.
  6. Classification. A dense layer with 100 outputs and a softmax produces the class probabilities.

Because the patch size must divide the image size evenly, 2×2 patches on a 32-pixel image produce exactly 256 patches. Changing the input resolution changes this count.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Configuration used in the example

The values below are the ones the example sets. They are a working configuration for this demonstration, not general recommendations for every dataset.

Setting Value in the example What it controls
Patch size 2×2 Size of each image patch
Patches per image 256 Sequence length fed to the encoder
Embedding dimension 64 Width of each patch representation
Attention heads 4 Number of attention heads per block
Transformer blocks 8 Depth of the encoder
Attention and projection dropout 0.2 Regularization inside each block
Batch size 128 Images per gradient update
Epochs 50 Full passes over the training set
Learning rate 0.001 Optimizer step size
Weight decay 0.0001 Penalty on weight magnitude
Label smoothing 0.1 Softens the one-hot targets in the cross-entropy loss

The example also holds out a validation split from the training data. The page describes this setting but the value is set in the code, so check the script before you copy it.

Cost: self-attention versus external attention

The Keras page gives an asymptotic account of the cost of each attention type. Here, d is the feature dimension, N is the number of patches, and S is the size of the external memory, which is a hyperparameter you choose.

Attention type Complexity stated on the page Dependence on sequence length
Self-attention O(d · N²) Quadratic in N
External attention O(d · S · N) Linear in N for fixed S

This is a theoretical scaling statement, not a measured runtime. The page does not report wall-clock timings or accuracy numbers for the two attention types, so do not read this table as a promise of speed or accuracy. The gain matters most when N is large; with only 256 patches, the practical difference on your hardware may be small.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Implementation outline

The example imports keras, layers, and ops, loads CIFAR-100, one-hot encodes the labels, and sets the input shape to 32, 32, 3. The simplified fragment below shows those parts; copy the full script from the official page for the complete model and training loop.

import keras
from keras import layers
from keras import ops

(x_train, y_train), (x_test, y_test) = keras.datasets.cifar100.load_data()
input_shape = (32, 32, 3)
num_classes = 100

The rest of the script defines the augmentation layers, the patch extraction and embedding layers, the external attention block, and the pooling and softmax head. Each block is built from two linear layers and two normalization layers around the shared memories, as described above.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Training setup

The example uses categorical cross-entropy with label smoothing, a weight decay term, a fixed learning rate, a validation split, a batch size of 128, and 50 epochs. Those values produce the example’s configuration. They are not tuned results, and the page does not report a final accuracy that you can compare your run against.

Running the example on your own setup

The page was created on 19 October 2021 and last modified on 18 July 2023. It does not pin a Keras version, so the code may need small changes on a current release. Before running it:

What’s actually slowing this PC down?

Pick the symptom - the matching free tool is one click away.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  • Install Keras and confirm which backend you are using.
  • Expect the first run to download CIFAR-100 automatically.
  • Run the full 50-epoch schedule on a GPU if you can; training on a CPU will take much longer.
  • If you switch datasets, update the input shape and class count together, and make sure the patch size still divides the image width and height.
  • If you change the image size, recalculate the patch count (for example, a 64×64 image with 2×2 patches gives 1,024 patches) and expect the cost to rise.

Keep the learning rate, weight decay, and label smoothing as starting points, and adjust them only after you have a baseline run on your own data.

The Bottom Line

“”

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.