High Accuracy Image to Text (OCR)
Extract text from an image with the largest OCR model — the most accurate, a bigger one-time download.
Drop images with text here or click to choose
Photos, scans and screenshots: JPG, PNG, WebP, GIF, BMP, TIFF or AVIF. Reads English and other Latin-alphabet languages, Chinese and Japanese. It can't read Korean. Files are processed in your browser and never uploaded. They stay in this browser until you remove them.
Ctrl+OchooseCtrl+Vpaste
Text recognition model
Accuracy model · …
Reads English and other Latin-alphabet languages (French, German, Spanish, Polish…), Chinese (simplified and traditional) and Japanese. It can't read Korean, Russian or other Cyrillic, Arabic, Thai, Hindi or Vietnamese.
Licenses: models (Apache-2.0) · engine (MIT)
When accuracy matters most — small print, photographed pages, receipts, low-contrast or busy images — this tool uses the largest of the three recognition models. The 132 MB model is downloaded once and kept in your browser. With graphics acceleration (WebGPU, in most current desktop browsers) a dense page is read in a few seconds; on the processor alone a short page takes about half a minute. Your images never leave your device, and it's free with no sign-up.
Compare the three Image to Text models
| Image to Text (Fastest) | Image to Text (Balanced) | Image to Text (Accuracy)This page | |
|---|---|---|---|
| Model download | 6 MB | 29.7 MB | 132.3 MB |
| First use: download with the engine at 50 Mbit/s | 19.6 MB · ≈ 3 sec | 43.4 MB · ≈ 7 sec | 145.9 MB · ≈ 24 sec |
| After the first use | Kept in your browser — no second download | ||
| Speed | Fastest | Fast | Slowest |
| Accuracy | Good on clear text | Better on photos and small print | Best on hard images |
| Reads | Latin alphabet (English, French, German…) and Chinese | Latin alphabet (English, French, German…), Chinese, and Japanese | Latin alphabet (English, French, German…), Chinese, and Japanese |
The recognition engine (13.7 MB, or 27.4 MB in browsers with WebGPU) is downloaded once and shared by all three. Nothing is downloaded until you ask the tool to read an image.
How to extract text with high accuracy
- Drop images anywhere on the page, paste one with Ctrl+V, or click to choose files. An image you already added in another image tool — say, on the Balanced page, to compare — is in the list already.
- The first time, click Download model & read text. The download (about 146 MB with the engine) happens once — the progress panel shows the size, the percentage and the time left — and it is then saved in your browser, so next time the selected image is read as soon as the page opens, and the others in the list with Read or Read all.
- Check the text beside the picture: click a region to find its line, double-click it to edit that line. Underlined lines are worth a second look. Then Copy text (Ctrl+Shift+C) or Download .txt.
- To share the picture without its text, download the masked copy under Hide the text in the image — pick the color and the regions to cover.
Features
- The largest model: 132 MB, about 146 MB with its engine, downloaded once and then saved in your browser
- Reads small, blurry, rotated or low-contrast text most accurately
- With WebGPU, a dense page in a few seconds; on the processor alone, about half a minute for a short page
- Reads English and other Latin-alphabet languages, Chinese (simplified and traditional) and Japanese
- Text in reading order, with the regions outlined and numbered on the image
- Edit, copy or download the text as .txt, and download a copy of the image with the text hidden
- Every downloaded file is checked before it is used, and files that arrived complete are kept if a download is interrupted
Is it private?
Yes. The model runs in your browser; the image and the recognized text are never uploaded, and the images stay in this browser until you remove them from the list. That makes it suitable for receipts, IDs, contracts and other documents you'd rather not send to an online service.
Frequently asked questions
How is this different from the other two tools?
It reads small, blurry, rotated or low-contrast text more correctly, because it uses the largest recognition model. The cost is a bigger one-time download (a 132 MB model, against 6 MB and 30 MB) and slower reading where the browser has no graphics acceleration: about half a minute for a short page and several minutes for a dense one, more on phones. For clear screenshots, Fastest or Balanced is enough.
Do I have to download it every time?
No. The model and its engine (14 MB, or 27 MB with WebGPU, shared with the other two tools) are downloaded once and saved in your browser, and every file is checked before it is used. Next time the tool starts with no download. Delete downloaded models frees the space whenever you like.
Which languages does it read?
Printed English and other Latin-alphabet languages (French, German, Spanish, Portuguese, Italian, Polish, Czech, Turkish…), digits and symbols, Chinese (simplified and traditional) and Japanese (kanji, hiragana and katakana). It can't read Korean, Russian or other Cyrillic, Arabic, Thai, Hindi or Vietnamese.
Can it read a page that is upside down or photographed at an angle?
Yes. An upside-down page is detected and read the right way round, and the lines of a slightly tilted photo are straightened before they are put in reading order, so they don't get mixed up. Vertical Chinese and Japanese text is read in columns from right to left.
What if the download is interrupted?
Just try again: the progress panel says what went wrong and offers Try again. Files that arrived complete and passed their check are kept, so trying again downloads only the rest; nothing half-downloaded is ever used.