A model does not look at your photograph. It looks at a few hundred summaries of small squares of it
🌐 इस लेख को हिन्दी में पढ़ें
In short: Vision-language models convert an image into a sequence of patch tokens before any reasoning happens. This guide explains patchification and position embeddings, why a fixed token budget means resolution is traded against detail, how tiling and thumbnails work around it, how a projector lets a language model read image tokens as if they were words, why the language prior then dominates and produces hallucinated objects and confidently misread charts, why text inside an image can override the picture, and what actually improves results.
Photograph a bill, paste it in, ask for the total. The model reads it correctly. Do the same with a denser bill — smaller type, more rows, a stamp across one column — and it returns a total that appears nowhere on the page, stated as plainly as the correct one was a minute earlier.
The instinct is to call this a reasoning failure. It usually is not. By the time any reasoning happens, the image has already been converted into a fixed number of tokens, and the digits you cared about may not have survived the conversion. As with text, the representation decided before the model saw anything sets the ceiling on what it can possibly do.
An image becomes a sequence
A transformer consumes sequences, so an image has to be made into one.
The standard method is almost crude. Cut the image into a grid of small squares — patches, typically fourteen or sixteen pixels a side. Flatten each patch into a list of numbers, and pass it through a learned projection that turns it into a vector of the same size the model uses for everything else. Because flattening destroys any sense of where the patch came from, a position embedding is added separately to record its place in the grid.
That is the whole trick. An image is now a sequence of a few hundred vectors, in the same mathematical space as text tokens, and the same architecture can process both. It is the reason a single model can accept a photograph and a question about it in one input.
The token budget is the story
Here is the consequence that explains most day-to-day behaviour. A fixed number of patches is a fixed information budget, and the number is not large — a few hundred tokens for an entire image, which is roughly what a couple of paragraphs of text costs.
So a large photograph has to be squeezed into that budget. Downscale it and anything finer than a patch — small print, a thin line on a chart, a handwritten digit — is averaged away before the model begins. That information is not merely deprioritised; it is absent, and no amount of careful prompting recovers it.
Current systems work around this with tiling: split a high-resolution image into several tiles, encode each one separately at full detail, and add a downscaled thumbnail of the whole so the model retains global context. It genuinely helps, and it costs tokens in direct proportion — which is why a detailed document image can consume as much of your context window as several pages of writing.
It also explains the single most useful practical habit with these models: cropping to the region you care about improves accuracy dramatically, using the same model and the same question. You have not made the model smarter. You have spent the entire budget on the part that mattered.
The model is not looking at your image. It is looking at a few hundred summaries of small squares of it, and anything finer than a square was gone before reasoning started.
How the picture and the words get joined
Two ideas do this work, and they are often confused.
The first is contrastive pretraining: train an image encoder and a text encoder together so that a picture and its caption land near each other in one shared space. That produces excellent matching between images and descriptions, which is what image search and the text-steering of image generators rely on.
The second is what current assistants actually use. A vision encoder produces patch vectors, and a small trained projector maps them into the exact space the language model uses for word tokens. The language model then reads them as though someone had typed them. This is elegant and has a consequence worth internalising: after the projection, the language model's priors are in charge. Where the visual evidence is weak, ambiguous or averaged away, the answer comes from what usually follows in text rather than from the picture.
What this predicts about the failures
Almost every characteristic error of these models falls out of the two facts above.
Hallucinated objects. Asked what is in a kitchen photograph, models will sometimes list a plausible item that is not there. This is the language prior completing a scene, not the vision encoder misfiring — and it is why the errors are always plausible rather than random.
Counting and precise spatial relations. Nothing in the pipeline maintains a persistent object with a location; there are patch tokens and attention. Asking whether the third item from the left is above or below the label is asking for something the representation holds only loosely.
Dense tables, handwriting, small print. These are budget failures. The information was compressed out.
Charts. The most quietly dangerous case. A model recognises the genre of a chart confidently and will narrate a trend that fits the genre, because chart-shaped images are overwhelmingly accompanied in training text by descriptions of trends. The narration sounds authoritative and the axis values may be invented.
Text inside the image overriding the image. Because the encoder has learned that written words are strongly informative, a label reading "banana" stuck onto an apple can flip the classification. In an assistant that can act, this is prompt injection arriving through the camera: instructions written on a photographed page are content the model reads and may follow.
What actually improves results
Crop and zoom before asking. Ask what is visible rather than for a conclusion — "list the values in column three" beats "what does this show". Where you have the underlying data, give the model the data rather than a picture of the data. For dense documents, run a purpose-built text and layout extraction step first and hand the model the text, which is the same hybrid conclusion that applies to search: use the specialised tool for the part it is good at, and the general model for the reasoning. And verify any number that matters, for the same reason you would verify a citation.
Why it matters for students and researchers
A great deal of Indian administrative and commercial life exists as images of documents: forms, bills, land records, handwritten registers, certificates photographed on a phone at an angle in poor light. This is exactly the workload where the token budget bites hardest, and it is made harder by scripts. Recognition of Devanagari, Tamil, Bengali and others in photographed documents is markedly weaker than for Latin script, for the same reason it is weaker in speech and in tokenisation — less training data, fitted after the fact. Anyone building document automation here will find that the interesting work is not prompting a model but deciding which parts of the pipeline should not be a model at all.
The transferable idea is the one this pairs with on the text side: every modality is discretised before a transformer sees it, and the discretisation sets the ceiling. Words become subword tokens, images become patch tokens, audio becomes frames. Each conversion throws something away — letters in one case, fine detail in another — and the resulting limits look like failures of intelligence while being failures of representation. Knowing which is which is most of what it takes to use these systems competently.
Frequently asked questions
How does an AI model read an image?
It cuts the image into small square patches, converts each into a vector, adds information about the patch's position, and processes the resulting sequence. The image becomes a few hundred tokens in the same space the model uses for text.
Why does a model misread a detailed document or small text?
Because the image is compressed into a fixed token budget. Detail finer than a patch is averaged away during downscaling, so it is absent before any reasoning begins and cannot be recovered by rephrasing the question.
Why does cropping an image improve the answer?
Because the whole token budget is then spent on the region you care about instead of being spread across the full frame, so the relevant detail survives the conversion at much higher effective resolution.
Why do models invent objects that are not in a photo?
Because after the image is projected into the language model's space, the language model's expectations fill gaps where visual evidence is weak. The invented item is whatever plausibly belongs in such a scene, which is why these errors always sound reasonable.
Can text written inside an image mislead a model?
Yes. Written words in an image are strongly informative to the encoder, so a misleading label can override the visual content — and for a system that can take actions, instructions written on a photographed page are a route for prompt injection.