Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Image-capable large language models (LLMs) do not read pictures as ordinary text. They process an image into a visual representation, combine that representation with your prompt, and generate a response. The details differ by model: one may divide an image into patches or tiles, another may resize it or apply other preprocessing. That is why resolution, image quality, and the exact question you ask can change what a model notices—and what it misses.

What happens when an LLM receives an image?

A useful way to think about image input is as a pipeline: the system prepares the image, encodes visual information, combines it with language input, and generates a text response. This describes the broad idea, not a universal implementation shared by every provider.

  1. Image input: The model or API receives an image in a supported form.
  2. Preprocessing: The service may resize, crop, tile, or otherwise prepare the image to fit its processing limits.
  3. Visual representation: A vision component encodes information from the image. Common approaches use patches or visual tokens, but their size and handling depend on the model.
  4. Multimodal processing: The system processes the visual representation together with your text prompt.
  5. Generated response: The language model produces an answer, description, classification, or other output based on those inputs.

A CVPR 2025 analysis describes an image encoder and adapter that produce image tokens. In the models analyzed in that paper, query-token representations carry global image information while details are extracted in a spatially localized way. That finding should not be treated as a complete account of every commercial model. Read the CVPR 2025 analysis.

Images are therefore not necessarily translated into a single caption before the model answers. The visual representation and prompt can be handled together, allowing a question such as “What does the sign say?” to direct attention toward text in the image. OpenAI’s GPT-4V system card describes an early vision-capable model and its evaluation; it is useful context, not a specification for all current systems. OpenAI’s GPT-4V system card.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How visual tokens, patches, and tiles work

Some vision systems divide an image into smaller regions so they can represent local details as well as the overall scene. Documentation uses different terms—such as patches, tiles, or image tokens—and providers do not share one standard image representation.

For example, Anthropic documents 28-by-28-pixel patches called visual tokens and model-tier limits on image size and token count. OpenAI documents model-dependent image detail modes, resizing behavior, patch budgets, and image-token accounting. Gemini documents tiling and a media-resolution control. These are provider-specific implementation details, not interchangeable rules or general measures of model intelligence. Check the relevant documentation for the model and API version you use: Claude vision, OpenAI image and vision guide, and Gemini image understanding.

Why image resolution matters

Higher resolution can preserve small print, fine lines, and subtle visual details. It can also increase the amount of image information processed, which may raise token use, latency, or computation. Resizing an image to fit a model’s limits can remove information, particularly when the original contains small text or dense detail.

Google’s Gemini image-understanding guide states: “Higher resolutions improve the model’s ability to read fine text or identify small details, but increase token usage and latency.” The precise tradeoff depends on the provider’s processing method, model, and settings; a resolution control or token rule from one API should not be assumed to apply to another.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

A 2026 ICLR paper on adaptive patching puts the other side of the tradeoff this way: “In principle, for general and straightforward multimodal understanding, low-resolution images are sufficient.” The paper also discusses how naive resizing can lose information and how high-resolution processing costs more computation. That is the paper’s framing of broad, straightforward tasks—not a guarantee for every image or model. Read the AdaPatch paper.

In practical terms, use an image with enough detail for the question. A broad scene-description request may not require the same source quality as reading a receipt or interpreting a chart. If an answer depends on a small region, crop that region or provide a clearer image when your tool permits it. Cropping can make the target easier to inspect, but it may remove context that matters to interpretation.

What image-capable LLMs can—and cannot—do

Depending on the model and application, image input can support captioning, visual question answering, classification, object detection, segmentation, and some OCR-like tasks. These labels describe task types, not a promise of exactness. An API may expose image understanding without guaranteeing specialized accuracy for every task.

  • Text in images: Models can often answer questions about legible text, but small, rotated, compressed, or non-Latin text can be difficult.
  • Charts and diagrams: A model may describe a chart yet misread values or struggle when meaning depends on color, line style, or a fine distinction in the legend.
  • Counting and location: Exact object counts and precise spatial localization can be unreliable, especially in crowded scenes.
  • Unusual perspectives: Panoramic and fisheye images can be harder to interpret than ordinary views.
  • Descriptions: A fluent answer can still contain a mistaken or invented detail.

OpenAI’s current vision guide explicitly warns: “Vision models can make mistakes.” Its documented limitations include small text, non-Latin text, rotated images, some charts, precise spatial localization, panoramic or fisheye images, and exact counting. Google’s guide describes common image-understanding tasks, while Anthropic recommends clear, legible input and attention to cropping, resizing, and compression artifacts. For all three providers, these guides explain implementation and usage rather than establishing a controlled cross-provider accuracy ranking. OpenAI vision guidance, Anthropic vision guidance, and Google Gemini image guidance.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

How to get a model to read text or inspect a detail

  1. Start with a clear source image. Avoid blur and compression artifacts; ensure the text is legible at the size supplied.
  2. Check orientation. Rotate the image upright before sending it, particularly if the text is sideways or upside down.
  3. Frame the relevant area. Crop around a receipt line, sign, or chart label if the full image makes it difficult to inspect. Keep enough surrounding context to preserve meaning.
  4. Ask a specific question. For example, ask the model to transcribe a particular sign or identify the value next to a named chart label rather than merely saying “read this.”
  5. Verify important results. Compare extracted text, counts, or measurements against the source image before using them in consequential decisions.

When results are poor, first distinguish an input problem from a model limitation. A clearer crop may help if the target is too small; it will not make an ambiguous chart or obscured text certain. Provider guides offer model-specific recommendations and limits, so consult the guide for your chosen API before changing image settings.

Using image understanding in an application

Developers should treat image handling as a provider-specific API contract. Confirm accepted image formats, size and resolution limits, how images are supplied, available detail controls, and how image input contributes to usage. Those settings can change and may differ between models within the same provider.

For an implementation comparison, useful questions include whether the API accepts your image format, whether it resizes or tiles images, what detail controls are available, how it reports image-token usage, and what it documents about latency or rejection behavior. OpenAI, Anthropic, and Google publish separate vision guides, but the cited documentation does not establish a controlled benchmark proving that one provider is more accurate than another. Choose based on the task, documented constraints, and your own validation with representative images.

Independent reader supportYour contribution helps us test, update, and keep practical guides available for everyone.Support on Ko-Fi

Or skip the browser setup

If your application first needs a screenshot of a web page to pass to an image-capable model, ScreenshotNeo is a website screenshot API and MCP server for developers. One GET request returns a PNG, JPEG, WebP, or PDF. Its clean-shot options can accept consent banners and remove more than 60 known consent platforms, newsletter popups, and chat widgets before capture; each step can be turned off. Bot checks and CAPTCHAs, blank pages, timeouts, failed loads, and cache hits cost nothing, and response headers report the page verdict and billing status. An MCP server provides take_screenshot, get_page_info, and capture_pdf tools for Claude, Cursor, and other MCP clients.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

For a single screenshot, supply your API key and target URL:

curl -G "https://api.screenshotneo.com/v1/shot" -d access_key=YOUR_API_KEY --data-urlencode url=https://stripe.com -o shot.webp

See the ScreenshotNeo API documentation for setup and options. One thousand screenshots per month are free with no card; paid plans start at $5 for 3,000. Sign up for the free plan.

Frequently asked questions

Does an LLM convert every image into a caption first?

No. Image-capable systems can combine a visual representation with your prompt; an intermediate caption is not a universal required step.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.

Does a larger image always produce a better answer?

No. More detail may help with fine print, but costs and provider limits can matter, and resizing or other preprocessing may alter the input. Match image detail to the task and verify important outputs.

Can an image model reliably count every object?

Not necessarily. Exact counting is a documented limitation for current vision models; check counts against the image when precision matters.

Are vision models from different providers directly comparable?

Not from the provider guides alone. They describe different implementations and settings, not a controlled cross-provider accuracy benchmark.

Product prices and availability are accurate as of the date/time indicated and are subject to change. Any price and availability information displayed on Amazon at the time of purchase will apply.

Special offer. See more information about Outbyte and uninstall instructions. Please review EULA and Privacy policy.