Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

iTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more

You can build a Vision Transformer (ViT) image classifier in Keras from scratch, and the official Keras example shows the full path. It trains on CIFAR-100 and, after 100 epochs, reports about 55% test accuracy and 82% test top-5 accuracy. Those numbers come from a from-scratch run on a small benchmark, and the example itself says they are not competitive results on CIFAR-100. Use the example as a clear map of the architecture and the code, not as evidence of what ViTs can reach.

How a ViT turns an image into a classification

A ViT does not scan an image with convolution filters. It cuts the image into a grid of patches, treats each patch as one token in a sequence, and lets a stack of Transformer blocks compare every patch with every other patch. The Keras example follows that pipeline in four stages.

  1. Patch extraction. The input image is split into equal, non-overlapping squares. The example resizes its CIFAR-100 images to 72 by 72 pixels and uses 6 by 6 patches, which gives 12 patches per side and 144 patches per image.
  2. Patch projection and position embedding. Each flattened patch is passed through a dense layer that maps it to a vector of length 64 (the projection dimension in the example). A learned position embedding is added to each vector so the model knows where the patch sat in the image. Without this, the model would see a bag of patches with no spatial order.
  3. Transformer blocks. Each block applies layer normalization, multi-head self-attention with 4 heads, a residual connection, and a small MLP with another residual connection. The example stacks 8 of these blocks.
  4. Representation and classification head. The final normalized outputs are turned into one vector and passed to a dense classifier that produces a score for each class. The step that turns patch outputs into one vector is a design choice, covered below.

What the official example sets up

The example was written by Khalid Salama and implements the ViT described in the original paper by Alexey Dosovitskiy and coauthors, An Image Is Worth 16×16 Words. The settings below are the tutorial’s own values. They are not defaults that suit every dataset or compute budget.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Setting Value in the Keras example Notes
Dataset CIFAR-100 50,000 training images and 10,000 test images
Input size 72 x 72 pixels Inputs are resized before patching
Patch size 6 x 6 pixels 144 patches per image
Projection (embedding) dimension 64 Length of each patch vector
Attention heads 4 Multi-head self-attention per block
Transformer layers 8 Blocks stacked in sequence
Epochs 10 as a test value; 100 for real training The example tells readers to switch to 100 for a real run

Because the example is a from-scratch model with no pretraining, the 10-epoch setting is only a quick check that the pipeline runs. Expect the 100-epoch setting to be the one that produces the published figures.

#1 Best Overall
Sale
Hands-On Machine Learning with Scikit-Learn, Keras, and TensorFlow: Concepts, Tools, and Techniques to Build Intelligent Systems
  • Use scikit-learn to track an example ML project end to end
  • Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
  • Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
  • Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
  • Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning

What the reported results do and do not show

From-scratch accuracy on CIFAR-100

The Keras page reports about 55% test accuracy and 82% test top-5 accuracy after 100 epochs. The page compares this with its own ResNet50V2-from-scratch figure of 67% accuracy and states that the ViT result is not competitive on CIFAR-100. Quote these numbers only with that configuration attached: one dataset, one from-scratch run, one set of tutorial hyperparameters, and the example’s 2021 page as the source.

Why pretraining matters

The original ViT paper reports its strongest transfer results after pretraining on JFT-300M, a large internal image dataset, and then fine-tuning on the target task. The Keras example names that pretraining step to explain the gap. A from-scratch ViT trained on a few tens of thousands of images does not have the same data advantage, which is why the example’s accuracy is modest.

Using your own labeled images

For a custom dataset, Keras’s image_dataset_from_directory utility builds a labeled dataset from a folder tree in which each subfolder is one class. The Keras from-scratch image-classification example shows JPEG loading from disk and a preprocessing and augmentation pipeline, which you can adapt to a ViT.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
  1. Arrange the folders. Create one subfolder per class, for example data/train/cats and data/train/dogs, and a matching data/val tree. Class labels come from subfolder names, so keep them consistent between train and validation.
  2. Load the datasets with the image size you chose. Set image_size to match the input size of your ViT. Your patch size must divide that size evenly, or the patch grid will not fit.
    import tensorflow as tf
    
    train_ds = tf.keras.utils.image_dataset_from_directory(
        "data/train",
        image_size=(72, 72),
        batch_size=32,
    )
    val_ds = tf.keras.utils.image_dataset_from_directory(
        "data/val",
        image_size=(72, 72),
        batch_size=32,
    )
  3. Set the class count from your data. The output layer of the classifier must have one unit per class. Set it from the number of subfolders, not from the CIFAR-100 value of 100.
  4. Add augmentation and normalization. Use Keras preprocessing layers such as random flips, rotations, or crops, applied to training data only. Scale pixel values the same way at training and inference time.

Folder labels, class count, image size, and augmentation are all dataset decisions. Tune them for your own images rather than copying the CIFAR-100 values.

Design choices you can change

Flattened outputs or a pooled representation

The original ViT paper prepends a learnable class token to the patch sequence and classifies from that token’s final state. The Keras example does not do this. It flattens the final Transformer outputs into one vector and passes that to the classifier. The page also names global average pooling, which averages the final patch outputs, as another valid way to aggregate them. Pick one and keep the classifier’s input size consistent with it.

Training from scratch or fine-tuning

Training from scratch is the path the example demonstrates. Fine-tuning a pretrained ViT is the path the paper’s strongest results depend on. If you have a modest labeled dataset and no pretrained weights, expect a from-scratch ViT to need careful tuning and may underperform a convolutional network. The Keras example itself uses ResNet50V2 as its comparison point.

Small-dataset variants

Keras also publishes a separate example that discusses shifted patch tokenization and locality self-attention for small datasets. These are distinct modifications, not the same model as the basic example. Start with the basic architecture to understand the pipeline, then consider the variant if your dataset is small.

Free tools Windows power users keep installed

One-click scans. No signup required.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Scope and version notes

  • Page date. The example page was created and last modified in 2021. Keras APIs and example code change over time, so check the current official page and your installed Keras version before copying code.
  • Hardware and runtime. The example does not give a hardware requirement or a runtime figure, and this article does not supply one. Training time depends on your GPU, batch size, and input size.
  • Reproduction. Results depend on the environment and random seeds. Expect your numbers to differ somewhat from the published ones.

The Keras example is a sound starting point for learning how a ViT works in code. It is also a clear reminder that transformer-based image models depend heavily on data scale and pretraining.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.