
Start here
Pointing a camera at a receipt, a chart, or a handwritten note and asking an assistant what it says is one of the more immediately useful things a multimodal model does. It is also one of the easiest places to be quietly misled, because a wrong description of an image tends to arrive in the same confident sentence structure as a correct one. The task is learning where that confidence outruns the underlying capability.
What the documents say
The clearest record of this comes from OpenAI's own GPT-4V(ision) system card, published 25 September 2023. It describes a real pilot with Be My Eyes, an organisation building tools for blind and low-vision users, run from March to August 2023 with a beta group that grew to 16,000 users. Testers surfaced hallucinations and 'errors, and limitations created by product design, policy, and the model'; one tester is quoted saying the tool 'very confidently told me there was an item on a menu that was in fact not there.' In a separate scientific-imagery test, the same system card reports that when two pieces of text sat close together in an image, the model would sometimes merge them into a term that does not exist. Anthropic's current vision documentation, retrieved as of 16 September 2026, states plainly that Claude 'might hallucinate or make mistakes when interpreting low-quality, rotated, or very small images,' that its counting and coordinate outputs are approximate, and that it 'is not designed to interpret complex diagnostic scans such as CTs or MRIs.' Google's Gemini vision documentation recommends clear, correctly rotated images for best results, without detailing specific error types the way the other two documents do.
Check this
For anything with a number that matters — a total on a receipt, a value on a chart — read the number back yourself rather than accepting the assistant's summary of it. Ask the assistant to quote the exact text it read rather than paraphrase it, since a paraphrase can smooth over a misread character in a way a direct quote cannot.
What holds and what fails
Vendors' own documentation and the Be My Eyes pilot agree on where this holds and fails: description and general context work well enough that a blind or low-vision user base of hundreds of thousands has relied on it, while Be My Eyes itself still warns its users not to use the tool for safety and health tasks like reading a prescription. It fails predictably on small or low-quality images, dense or closely spaced text, and any task requiring precise counting or diagnostic interpretation.
- Read back any number that has a consequence, rather than trusting a summary of it.
- Ask for an exact quote of text in an image, not a paraphrase.
- Do not use image understanding for medical, safety, or diagnostic decisions.
The same confident tone describes both a correct reading and a fabricated one, so the check has to happen on the number or the text, not on how sure the answer sounds.
Sources & reading trail
Documents the Be My Eyes pilot in which blind and low-vision beta testers found hallucinations and confidently stated errors, and a scientific-imagery test where the model merged separately located text.
Source published: 25 September 2023 · Retrieved: 16 September 2026
States current limitations of Claude's image understanding: hallucination on low-quality or small images, approximate counting and spatial coordinates, and unreliability for diagnostic medical scans.
Source published: Not established · Retrieved: 16 September 2026
States Gemini's supported vision tasks and recommends verifying image orientation and clarity for best results.
Source published: Not established · Retrieved: 16 September 2026
Documentation, regulator guidance and studies establish the record; the checks and the boundary are AI Use Field Guide editorial analysis. This retrospective draft does not imply the site published on the event date.