Describe any image with AI

Upload a photo and get a plain-English description of what is in it — plus every line of text the image contains. One upload, both answers.

This only affects the text reading. The description works the same whatever the picture is.

?

Drop an image to describe it

JPG, PNG or WEBP — up to 20 MB. Each image uses one of your monthly requests. Your images are visible only to your account and auto-deleted after 30 days.

How it works

  1. Upload one image. Drag it in or browse for it.
  2. Two models read it. A vision model writes the description; the OCR engine reads any text, exactly as the Image To Text tool does.
  3. Read both. The description first, the transcribed lines below it.

Why use it

Description and text together

Most tools give you one or the other. This gives what the picture shows and what it says, from a single upload.

Nothing gets refused

No content filter sits in front of the models — they read the image you give them.

Free, 500 a month

Sign in with Google and describe up to 500 images every month at no cost.

Private by default

Your uploads and results are visible only to the account that created them, and are deleted after 30 days.

What it gets right, and where it does not

A caption is a guess. Knowing which guesses hold up saves you from trusting the wrong one.

Scenes and objects: reliable

What kind of place, what is in the frame, the dominant colours, roughly how things are arranged. This is what the model is trained for and it is usually right.

Counting: unreliable

Ask it how many of something there are and it will answer with confidence whether or not it counted. Treat any number in the description as an estimate.

Text: read the OCR, not the caption

The description may mention that there is writing, and may even paraphrase it wrongly. The transcribed lines below the description are the ones actually read glyph by glyph.

Named people and places: no

It describes what a thing looks like, not who or where it is. It does not recognise individuals, and any name it produces is a guess you should discard.

Frequently asked questions

1. What does it actually give me?

Two things from one upload: a description of what the image shows — one short line and one paragraph — and every line of text found in the image, transcribed.

2. How is this different from Image To Text?

Image To Text reads the glyphs and nothing else. This tool does that too, and adds a description of the picture itself. If all you need is the text, the other tool is faster.

3. Can I trust the description?

Treat it as a useful first pass, not as a fact. It is written by a vision model that produces fluent sentences whether or not it is sure, so it can be wrong without sounding wrong. Counts and any proper nouns are the least reliable parts.

4. Which model writes the description?

Moondream 2, an open vision-language model under the Apache 2.0 licence. The exact version is shown with every result, because a description is only interpretable if you know what produced it.

5. Does it refuse any images?

No. There is no content filter anywhere in the pipeline. What it will not do is identify people: it describes appearances, not identities.

6. Is it really free?

Yes. Signing in with a Google account gives you up to 500 requests per month, free.