Do these 3 things before closing this tab:
1Clear out junk files and repair common Windows errors2Fix the driver behind crashes, sound loss and screen glitches3Repair Windows errors before they cause bigger problemsiTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
A Keras dual encoder can retrieve images from ordinary text by learning a shared embedding space: one model encodes images, another encodes text, and a search ranks indexed image vectors against the query vector. Khalid Salama’s Keras example demonstrates this approach with Xception, BERT, and MS-COCO. It is an illustrative implementation from January 30, 2021—not a guarantee of present-day package compatibility or search quality.
What a dual encoder does
A dual encoder, also called a two-tower model, has separate image and text encoders. Training brings representations of matching captions and images closer together in a shared vector space. At search time, the system encodes a natural-language query and compares it directly with image representations computed in advance.
The Keras example is inspired by CLIP. Its training objective uses caption-image dot-product similarities and cross-entropy; the target similarities also account for caption-caption and image-image similarities. Once training is complete, the example uses the separately fine-tuned vision and text encoders for retrieval and discards the combined training model.
Free tools Windows power users keep installed
One-click scans. No signup required.
Models and data in the Keras example
Image and text towers
The vision tower uses ImageNet-pretrained Xception without its classification head and with average pooling. Its input is a 299-by-299 RGB image, processed with Xception preprocessing, then passed through projection layers. The text tower uses an uncased small BERT model and its preprocessing from TensorFlow Hub; a projection head maps the pooled BERT output to the same dimensionality as the image representation. The base encoders are frozen by default in this example.
#1 Best Overall
- Use scikit-learn to track an example ML project end to end
- Explore several models, including support vector machines, decision trees, random forests, and ensemble methods
- Exploit unsupervised learning techniques such as dimensionality reduction, clustering, and anomaly detection
- Dive into neural net architectures, including convolutional nets, recurrent nets, generative adversarial networks, autoencoders, diffusion models, and transformers
- Use TensorFlow and Keras to build and train neural nets for computer vision, natural language processing, generative models, and deep reinforcement learning
MS-COCO training sample
The tutorial describes MS-COCO as containing more than 82,000 images, each with at least five caption annotations. Its configuration samples 30,000 training images and two captions per image, yielding 60,000 caption-image pairs. The tutorial also reports a 13 GB compressed image archive; that is a figure for the archive it describes, not a general storage estimate for other image collections.
How retrieval works
- Index images: Run the vision encoder over the collection and store each image embedding together with its path or other identifier.
- Encode the query: Pass the user’s natural-language description through the text encoder.
- Score candidates: L2-normalize the query and image embeddings in the example retrieval function, then compute dot products between the query vector and image vectors.
- Return results: Select the top-k scores, map their indices back to image paths, and display those images.
The tutorial’s example queries include “a plate of healthy food,” “a woman wearing a hat is walking down a sidewalk,” and “a bird sits near to the water.” Its worked example searches for “a family standing next to the ocean on a sandy beach with a surf board.” These illustrate the kind of descriptive text the model accepts; they are not evidence about typical user queries.
Rank #2
Exact matching and larger collections
The demonstration computes exact dot-product matches against the image vectors. That is straightforward for a modest collection, but scanning every vector for every query may not suit a large real-time service. The tutorial names ScaNN, Annoy, and Faiss as approximate similarity-search options for larger collections. It does not benchmark or rank them, so the choice should be validated against the application’s latency and retrieval-quality requirements.
The Tool Desk
Outbyte Driver Updater FREEFix the driver behind crashes, sound loss and screen glitchesFind Drivers →Outbyte PC Repair FREERepair Windows errors before they cause bigger problemsFix Now →For bulk generation of image embeddings, the page also mentions Apache Spark and Apache Beam as possible parallel processing frameworks. These are implementation options rather than measured recommendations. Embedding cost and throughput, how often images change, and the time needed to refresh the index are operational factors to assess for a production collection.
Rank #3
What the reported result means
The tutorial reports 6.235% evaluation top-k accuracy for its described run. Its evaluation checks whether the associated image appears among the top 100 results for captions against out-of-training-sample images. This is a result from that tutorial’s dataset sample, model setup, and evaluation—not a general benchmark, nor a prediction for a different collection, top-k value, or query set.
The page’s timing figures are similarly specific to its run: its prose estimates around 12 minutes per epoch with a V100 GPU and around 8 minutes with two GPUs for 60,000 pairs at batch size 256, while displayed output records about 9 minutes per epoch on two GPUs. These historical, hardware- and run-specific figures should not be treated as current training-cost estimates.
Rank #4
Ways to improve a trained search system
Salama’s tutorial suggests increasing the training sample and number of epochs, trying other image and text backbones, making the base encoders trainable, and tuning hyperparameters—especially the loss temperature. These are proposed avenues to test, not improvements demonstrated as comparative results on the page.
Recommended Free Tools
For a meaningful evaluation, use held-out queries and images representative of the intended collection, and state the retrieval metric and top-k threshold. Also compare exact and approximate retrieval at the collection size and latency that matter to the application. Measure embedding-generation throughput and index-refresh needs separately; the tutorial does not provide production measurements for those operational factors.
Best Value
- THE FASTEST WAY TO PHONICS MASTERY - Teach and Learn Phonics with Audio Sounds, learners get to see the spelling pattern and hear the related phonetic sounds. The audio reinforcement demonstrates the content and solidifies the learning quicker than flash cards and workbooks.
- PHONICS SYSTEM QUIZZES THEM IN 13 STEPS - The electronic phonics workbook starts with single letter sounds like a, b and c. This progresses through short and long vowel sounds, consonant digraphs, trigraphs, diphthongs, bossy R, silent letters and irregular phonics.
- TEST AND BUILD PHONEMIC AWARENESS - Our Educational Learn to Read Machine challenges them to find words which contain a particular phonetic sound or pick out phonetic sounds from the given vocabulary. All created with American English Audio.
- LEARNING THAT CHILDREN ENJOY - The Screenless Educational Tablet With Talking Flash Cards tests and quizzes children on their reading and phonics knowledge while correcting errors and compounding knowledge, all the while putting a smile on their face.
- UNLOCK YOUR CHILD'S POTENTIAL WITH BAMBINO TREE! - From numbers and pictures bingo to letter flashcards and phonics games, we offer a variety of learning materials and games for children with effective tested teaching strategies.
Compatibility and implementation status
The Keras tutorial was created and last modified on January 30, 2021. Its setup specifies TensorFlow 2.4 or higher and dependencies from that period, including TensorFlow Hub, TensorFlow Text, and TensorFlow Addons. Check current package compatibility before following its installation steps rather than assuming those instructions describe a modern environment. The original implementation and code are available in the Keras natural-language image search tutorial.
The project’s Hugging Face model card says loading through its TF-Keras path requires keras<3.x or tf_keras, and describes a reproduction trained on 30,000 images. It also states that the model is not deployed by an Inference Provider; that is the card’s reported status, not a claim that independent third parties cannot run it.
Quick Recap
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
What’s actually slowing this PC down?
Pick the symptom - the matching free tool is one click away.

