Do these 3 things before closing this tab:
1Scan for outdated or missing drivers - takes under a minute2Repair Windows errors before they cause bigger problems3Fix the driver behind crashes, sound loss and screen glitchesiTechGuides is reader-supported. When you buy through links on our site, we may earn an affiliate commission. As an Amazon Associate I earn from qualifying purchases. Learn more
Gemini API File Search can ground answers in your own indexed files, but “multimodal” has a narrow meaning here. File Search indexes text and images. It does not index audio or video, even though other Gemini input routes may accept media. To search images, you must override the store’s default text-only embedding model and configure models/gemini-embedding-2. Accepted images are PNG or JPEG files no larger than 4K × 4K pixels.
What “multimodal” means in File Search
Gemini models accept several kinds of input, but a File Search store accepts a smaller set. The table below separates what the store can index from what other Gemini input methods might handle. Check the File Search documentation for the current list before you design around a format.
| Content type | Indexed by File Search? | Embedding model | Constraints |
|---|---|---|---|
| Text documents | Yes | gemini-embedding-001 (the default text setup) |
File type and size limits for text not stated on the File Search page |
| Images | Yes | models/gemini-embedding-2, configured as an override of the text-only default |
PNG or JPEG only; maximum 4K × 4K pixels |
| Audio | No | Not applicable | Not currently supported by File Search |
| Video | No | Not applicable | Not currently supported by File Search |
If your corpus includes recordings or video, File Search will not index them directly. You would need to convert that content into text or still images before import, and File Search does not do that conversion for you.
Free tools Windows power users keep installed
One-click scans. No signup required.
How File Search retrieves content
File Search is a managed retrieval-augmented generation workflow. It imports your files, splits them into chunks, and indexes those chunks. When a request arrives, the system embeds both the query and the imported content, finds similar and relevant chunks, and passes those chunks to the model as context for the answer. The model therefore answers from the retrieved material, not from the entire file collection.
This design has a practical consequence. Retrieval quality depends on how well the chunks and embeddings represent your content. The documentation does not publish accuracy or retrieval-quality benchmarks comparing text and image retrieval, so measure results on your own corpus before relying on them.
Build the pipeline
- Create a File Search store. For text-only retrieval, use the documented default text embedding,
gemini-embedding-001. For image retrieval, configure the store to usemodels/gemini-embedding-2in place of the text-only default. - Upload or import files into the store. Use one of the documented upload or import workflows. For images, confirm each file is PNG or JPEG and no larger than 4K × 4K pixels before you send it. Resize anything larger first.
- Poll the operation until it finishes. If the method you chose returns a long-running operation, wait for its completion status before querying the store.
- Send a Gemini request with the File Search tool pointed at your store. The documentation shows both
generateContent-style examples and newer Interactions examples. Confirm which API surface your project uses, and use the current SDK reference for exact syntax, since SDKs change. The page includes Python, JavaScript, Java, and REST examples. - Inspect the annotations in the response. Text citations identify the source file. Image citations can include a
media_id, which you can use to download the referenced image chunk.
Reading citations and tracing answers
Citations let you see which uploaded material supported an answer. For text, the citation points to the source file. For images, the media_id lets you retrieve the exact image chunk the model used. Use these annotations to trace sources, but still check generated conclusions against the original material. A citation shows what was retrieved, not that the model interpreted it correctly.
Rank #2
File Search or direct file input?
File Search is the persistent, indexed path. You build a store once and retrieve from it across many requests. Supplying a file directly in a request is a separate path. Google’s file input methods guide says the right choice depends on file size, where the data is stored, and how often it will be used, and it lists availability across Batch, Interactions, and Live API endpoints.
| Comparison point | File Search | Direct file input |
|---|---|---|
| Pattern | Persistent indexed store, retrieval across many files | File supplied as part of a request |
| Modalities | Text and images (PNG, JPEG); no audio or video | Not stated in the file input guide’s comparison with File Search |
| Size and resolution limits | Images up to 4K × 4K pixels | The guide’s local PDF example uses a 50 MB limit; this does not apply to every file method or format |
| Endpoint and SDK compatibility | Shown in generateContent-style and Interactions examples |
Batch, Interactions, and Live API endpoints listed in the guide |
| Source and media citations | Text file citations; image media_id values |
Not stated in the file input guide |
| Retention and deletion | Indexed store data persists until manually deleted or the model is deprecated | Not stated in the file input guide |
| Cost | Embedding charged at first indexing; query-time embedding and storage free; normal token charges apply | Not stated in the file input guide |
Use File Search when your application repeatedly retrieves from a collection. Use direct input when a single file is part of one request and you do not need a reusable index. Confirm the file size limit and endpoint for your method before you commit to either approach.
Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Retention and cost
The File Search documentation separates temporary upload objects from the indexed store. Its current statements are:
- Raw File API objects are deleted after 48 hours.
- Indexed store data persists until you delete it manually or the model is deprecated.
- Embedding generation is charged when files are first indexed.
- Embedding generation at query time and File Search storage are described as free.
- Normal Gemini model input and output token charges apply to requests.
These are the documentation’s billing statements as of October 2026, not an estimate for any workload. Google’s page does not show a publication date, so check current pricing in the File Search documentation before you budget.
Quick Recap
Best Value
Rank #4
Checks before you ship
- Confirm the store’s embedding model. Text-only indexing uses
gemini-embedding-001. Image indexing requires themodels/gemini-embedding-2override. - Verify each image’s format and dimensions before upload. Only PNG and JPEG files up to 4K × 4K pixels are documented as supported.
- Wait for every import or upload operation to finish before running production queries.
- Match your request code to the API surface and SDK version you use, since the examples include two styles.
- Record how you will delete the store, because indexed data persists until you remove it.
- Test retrieval on representative questions and images from your own corpus.
Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.
Recommended Free Tools

