Toolonit
  1. Home
  2. Image
  3. Image to Text (Balanced)

Image to Text (OCR)

Extract text from an image with a mid-size OCR model — a balance of speed and accuracy.

Drop images with text here or click to choose

Photos, scans and screenshots: JPG, PNG, WebP, GIF, BMP, TIFF or AVIF. Reads English and other Latin-alphabet languages, Chinese and Japanese. It can't read Korean. Files are processed in your browser and never uploaded. They stay in this browser until you remove them.

Ctrl+OchooseCtrl+Vpaste

Text recognition model

Balanced model · …

Reads English and other Latin-alphabet languages (French, German, Spanish, Polish…), Chinese (simplified and traditional) and Japanese. It can't read Korean, Russian or other Cyrillic, Arabic, Thai, Hindi or Vietnamese.

 

Licenses: models (Apache-2.0) · engine (MIT)

Copy the text out of a photo of a document, a scanned page or a screenshot, then edit it and save it. This is the everyday choice of the three Image to Text tools: its 30 MB recognition model reads small print, photos taken at an angle and busy backgrounds noticeably better than the fastest one, and still reads a screenshot in seconds. It all runs in your browser: your images never leave your device, and it's free with no sign-up.

Compare the three Image to Text models

Image to Text (Fastest)Image to Text (Balanced)This pageImage to Text (Accuracy)
Model download6 MB29.7 MB132.3 MB
First use: download with the engine at 50 Mbit/s19.6 MB · ≈ 3 sec43.4 MB · ≈ 7 sec145.9 MB · ≈ 24 sec
After the first useKept in your browser — no second download
SpeedFastestFastSlowest
AccuracyGood on clear textBetter on photos and small printBest on hard images
ReadsLatin alphabet (English, French, German…) and ChineseLatin alphabet (English, French, German…), Chinese, and JapaneseLatin alphabet (English, French, German…), Chinese, and Japanese

The recognition engine (13.7 MB, or 27.4 MB in browsers with WebGPU) is downloaded once and shared by all three. Nothing is downloaded until you ask the tool to read an image.

How to extract text from an image

  1. Drop images anywhere on the page, paste one with Ctrl+V, or click to choose files — several at once is fine. Images already added in any image tool, the Fastest and Accuracy pages included, are waiting in the list.
  2. The first time, click Download model & read text. The model (30 MB) and the engine that runs it — about 43 MB together — are downloaded once and kept in your browser, so later an image you add is read at once, the selected one as soon as you open the page, and the others already in the list with Read or Read all.
  3. Compare the text with the picture: click a region to find its line, or a line to find its region. Underlined lines are the ones the model is unsure about. Edit, then Copy text (Ctrl+Shift+C) or Download .txt.
  4. To share the image without its text, open Hide the text in the image, pick a color, leave visible any region you want to keep, and download the masked copy.

Features

  • Reads printed English and other Latin-alphabet languages, Chinese (simplified and traditional) and Japanese
  • Text in reading order: columns one after another, a receipt's items next to their prices
  • Text regions outlined and numbered on the image; click a region to find its line
  • Edit the text, then copy it or download it as .txt; lines the model is unsure about are underlined
  • Hide the text: download a copy of the image with the text covered by boxes
  • Several images at once; upside-down pages are detected and read the right way up
  • One download of about 43 MB (the 30 MB model and its engine), then saved in your browser

Is it private?

Yes. The image is read by a model running in your browser; neither the picture nor the text is sent anywhere, and your images stay in this browser until you remove them from the list. Only the model and its engine are downloaded, once, from this site, and they stay saved in your browser.

Frequently asked questions

Which languages does it read?

Printed English and other Latin-alphabet languages (French, German, Spanish, Portuguese, Italian, Polish, Czech, Turkish…), digits and symbols, Chinese (simplified and traditional) and Japanese (kanji, hiragana and katakana). It can't read Korean, Russian or other Cyrillic, Arabic, Thai, Hindi or Vietnamese.

Why is the first run slower?

Because only the first time, the browser downloads the model and the engine that runs it (14 MB, or 27 MB in browsers that use the graphics card through WebGPU) and prepares them. After that they are saved, so reading starts with no download, and a page that is already open keeps working without a connection. If you already used another Image to Text tool, the engine is there and only the model is downloaded.

Fastest, Balanced or Accuracy — which should I use?

Balanced suits most images. Fastest (a 6 MB model, about 20 MB to download the first time) is enough for clear screenshots and documents in English or Chinese. Balanced (30 MB, about 43 MB) is more accurate on photos and small text and also reads Japanese. Accuracy (132 MB, about 146 MB) reads the hardest images best, but without graphics acceleration (WebGPU) it needs about half a minute for a short page. Each is downloaded once and then saved in your browser.

Does it keep the layout of the page?

It keeps the reading order, not the exact layout: pieces of text on the same row are joined, each row is a line, and a blank line separates paragraphs. Side-by-side columns are read one after the other, and tables come out as text in reading order, not as a spreadsheet.

Can I read several images at once?

Yes. Add as many as you like; each is read in turn. Copy all images copies every text under its file name, and Download all (ZIP) saves the masked copies.

Related tools