Extract text from any image

Batch up to 10 images — OCR each, then copy or download all at once.

Drop up to 10 images to extract text

JPG, PNG or WEBP — up to 20 MB each, 10 images per batch; each image uses one of your monthly extractions. Your images are visible only to your account and auto-deleted after 30 days.

How it works

From image to editable text in three steps.

1

Upload

Drag & drop or browse a JPG, PNG, or WEBP file — or paste an image URL, or shoot one with your camera.

2

Pick source language

Choose the source language — Chinese, Japanese, or one of 40+ Latin-script languages. English inside Chinese images is read too.

3

Copy text

Grab clean, editable text with the exact position of every line — copy one line or all of it.

Why use it

Positions included

Every line comes back with its exact box on the image, highlighted as you hover.

Chinese & Japanese ready

Tuned OCR for CJK glyphs — no more garbled kanji or hanzi.

Free, 500 a month

Sign in with Google and extract up to 500 images every month at no cost.

Private by default

Your uploads and results are visible only to the account that created them.

Getting good results from a photo

What this tool handles well, where it struggles, and how to read what it gives back.

Resolution decides almost everything

OCR quality is bounded by pixels per character, and that single factor outweighs every other setting. Before detection runs, an image is fitted onto a square canvas — so a 3024×4032 phone photo is scaled down, and body text that was 24 pixels tall in the original arrives at the model much smaller. When we raised that canvas from 800 to 1600 pixels, whole passages that had been coming back unread from photographs started reading correctly. The practical version for you: fill the frame with the text you care about. One panel of a menu photographed close beats the whole menu photographed from a step back, every time.

What reads well, and what does not

  • Reads well: printed text at a straight angle, even lighting, screenshots and scans, signage and packaging, menus shot close, Latin-script documents.
  • Struggles: handwriting (not supported), heavy motion or focus blur, a hand or phone shadow falling across the text, brush-style and highly decorative fonts, text smaller than roughly 20 pixels tall in the original file.
  • Partially supported: vertical Japanese and Chinese. Columns are reassembled into reading order (top to bottom, then right to left), but recognition on degraded vertical text is still the weakest path — see the detailed write-up.

Reading the coverage figure

Results show how many text boxes were detected and how many were actually read. That ratio — coverage — is more useful than an accuracy percentage, because a tool that discards uncertain lines can report a high average confidence precisely because it threw away the hard ones. If coverage is low, the model found text it could not read: usually a resolution or focus problem, and reshooting closer fixes it. If coverage is high but a line you can see is missing from the list, the detector never found it at all — which is a different failure, and the next section is the answer to it.

When a line is missing, tap it

Text the detector skipped produces no entry and no score, so nothing marks its absence. On the result image you can tap the missed text directly: the area is located, shown as an adjustable box, and read on confirmation. It does not count against your monthly quota, since it is our miss you are correcting. Lines read this way have no confidence threshold applied — you asked for that region specifically — so anything the model was unsure of is marked unsure in the list, and worth checking against the image.

What the output contains

Every line comes back with its text, a confidence score, and the four corner coordinates of its box on the image. Hovering a line highlights its box; clicking copies it. The JSON download contains all three for every line, which is what makes the output usable as input to something else — sorting by position, filtering by score, or cropping regions programmatically. Up to 10 images run in one batch, each as its own job, so one failure does not sink the rest.

Languages and models

Chinese (with English inside the same image) and Japanese run on the PP-OCR pipeline we operate ourselves, with per-language score thresholds tuned for photographs rather than clean scans. Forty-plus Latin-script languages run on a separate multilingual model, kept in isolated code so changes there cannot affect Chinese and Japanese quality. Korean, Cyrillic, Arabic and Thai are not supported as source languages today.

Your images

Uploaded images are processed and then deleted immediately — not retained for training or quality review. Results stay in your account so you can reopen them, and the Privacy Policy sets out exactly what is stored and for how long.

Frequently asked questions

1. Which languages can it read?

Chinese (Simplified) together with English, Japanese, and 40+ Latin-script languages including French, German, Spanish, Portuguese, Italian, Dutch, and Vietnamese. Pick the source language before uploading for best accuracy.

2. Does it give me where the text is?

Yes. Every extracted line includes its position on the image, drawn as a box over the photo. Hover a line in the list to highlight its box.

3. What file types can I upload?

JPG, PNG, and WEBP images up to 20 MB.

4. Is it really free?

Yes. Signing in with a Google account gives you up to 500 extractions per month, free.

5. Are my uploaded images stored?

Images are stored privately under your account so you can revisit recent extractions, and are automatically deleted after 30 days. No one else can access them.