Skip to content
TabBench

Image to Text (OCR)

Extract text from screenshots, scans and photos. Runs on your device, with an optional AI mode for handwriting.

Runs on your device by default. The optional cloud AI mode sends your input to Google Gemini.

What the Image to Text (OCR) does

OCR turns a picture of text back into text you can select, search and edit. It matters more than it sounds: a scanned contract, a screenshot of an error message, a photographed receipt — all of them hold information your computer cannot read, because to the machine they are just coloured pixels arranged in shapes. This tool offers two recognisers with genuinely different trade-offs. The default runs entirely inside your browser: the image is never uploaded, it works offline once loaded, and there is no usage limit. The optional AI mode sends the image to Google and reads things the on-device engine cannot — handwriting, table layouts, and scripts other than Latin. Which one you should use depends less on quality than on what is in the picture.

How to extract text from an image

  1. Upload a PNG, JPG, WebP or BMP, or paste a screenshot straight from your clipboard.
  2. Leave the recogniser on On-device unless you need what the AI mode adds. On-device keeps the image on your machine.
  3. Press Extract text. The first on-device run downloads about 9MB of recognition data; after that it is cached and near-instant.
  4. Read the confidence score. Below about 70% you should expect mistakes and check the result against the image.
  5. Correct anything wrong directly in the output box, then copy it or download it as a .txt file.

The Image to Text (OCR) runs on your device by default. If you switch to the optional cloud AI mode, your input is sent to Google's Gemini model to produce the result.

When to use it

Getting text out of a scanned PDF

A scan has no text layer, which is why converting one to Word produces an empty document and why a PDF editor cannot find any words to change. Export the page as an image, run it through here, and you have text again. This is the single most common reason people need OCR.

Copying from a screenshot

Error messages, chat threads and slides are constantly shared as images. Rather than retyping a stack trace by hand, extract it and paste it into your terminal or a search box.

Digitising receipts and invoices

Photographs of receipts are awkward: the paper curves, the lighting is uneven, and the layout is columnar. The AI mode handles all three considerably better than the on-device engine, though it means uploading the image.

Reading handwriting

Classical OCR is built around printed letterforms and does poorly on handwriting. If your image is handwritten, the on-device mode will likely return nonsense and the AI mode is the only realistic option.

Good to know

  • Resolution matters more than file size. A sharp 1000px-wide crop of the text beats a 12MP photo of the whole page.
  • Straighten the image first if it was photographed at an angle: on-device OCR assumes roughly horizontal lines of text.
  • Crop to just the region you need. Less surrounding clutter means fewer spurious characters.
  • Low contrast is the most common cause of poor results. Dark text on a light background reads far better than grey on grey.
  • The on-device engine is English-only here. For other scripts, use the AI mode.

Frequently asked questions

Is my image uploaded anywhere?

Not in the default on-device mode: recognition runs in your browser and the image never leaves your device. The optional AI mode does upload it to Google's Gemini API, and the tool says so before you use it.

Why is the first run slow?

On-device mode downloads a recognition model of about 9MB the first time. Your browser caches it, so later runs start immediately and work offline.

Can it read handwriting?

On-device OCR is poor at handwriting. The AI mode handles it well, along with tables and non-Latin scripts.

Which mode should I use for something confidential?

On-device, without exception. It runs in your browser and the image never leaves your machine, so an ID card, a bank statement or a medical letter stays with you. The AI mode uploads the image to Google and should not be used for anything you would not email.

Why is the accuracy lower than my phone's built-in scanner?

Phone scanners pre-process aggressively — deskewing, sharpening and thresholding the image before recognition, often using a dedicated model. Here you get the raw recogniser, so preparing the image yourself (crop, straighten, increase contrast) makes a large difference.

Does it keep the original layout?

Reading order is preserved and line breaks are usually right, but columns and tables are flattened by the on-device engine. The AI mode is asked to keep table rows together with tab-separated columns, which holds up reasonably well.