How to extract text from a scan, offline
OCR turns a picture of a page into text you can select, edit and paste. Where that conversion happens is the part worth asking about, because the input is the whole document — the salary figure, the diagnosis, the account number.
This guide covers what "on-device" means in practice, which scripts ScanCrypt recognises, and the exact steps to get the text out of a scan.
On-device OCR versus OCR in someone else's data centre
Recognition is expensive to build and cheap to rent. That is why a large share of scanner apps do it by uploading the page image to a server, running OCR there and sending the text back. The feature works, it is often more accurate on hard inputs, and it means a full-resolution copy of your document sat on a machine you do not control.
On-device OCR runs the recognition model on your phone. The page never becomes an HTTP request. There is nothing on the other end to be retained, indexed, subpoenaed or breached, because there is no other end.
The distinction is invisible in the interface. Both routes show you a spinner and then some text. The test that tells them apart is airplane mode: turn the radio off and run OCR. If the app returns text, the model is on your phone.
- Uploaded OCR sends the page image, not just the text it finds.
- A privacy policy that permits "processing" of your files is describing an upload.
- On-device OCR is testable from the outside — no source code needed.
Which scripts ScanCrypt reads
ScanCrypt ships five recognisers and lets you choose between them on the OCR screen, in a row labelled SCRIPT. The choice matters: a recogniser trained on one script will not read another, so a Hindi page run through the Latin model comes back as noise rather than as an error message.
Pick the script before you run the extraction. Latin is the default and covers English and the European languages that use the Latin alphabet.
- Latin — the default, for English and European languages.
- Chinese, simplified and traditional.
- Devanagari, which covers Hindi, Marathi and Sanskrit.
- Japanese.
- Korean.
What you can do with the text
The result lands in an editable field, not a read-only panel. OCR on a creased receipt or a faint carbon copy will get some characters wrong, and fixing them in place before you use the text is faster than fixing them at the destination.
From the OCR screen the text can be copied to the clipboard, shared to any app through the standard Android share sheet, or saved as a .txt file. All three act on whatever is in the field at that moment, including your corrections.
- Copy — for pasting into a form, a message or a spreadsheet.
- Share — hands the text to any app that accepts text.
- Save as TXT — writes a plain text file you keep.
Which tier OCR is on
OCR is a Pro tool in the current release. Opening it on the free tier shows the upgrade screen instead of the extraction. Pro is a one-time purchase or an annual subscription, and it also removes the ads.
Scanning itself is free, so you can capture the document first and decide about OCR afterwards.
How to do it in ScanCrypt
- Get the document into ScanCrypt. Scan it, or import a PDF or image. OCR runs on a file in your library.
- Open OCR. From Tools, tap OCR and pick the file. From a document you already have open, use the OCR action in the viewer.
- Pick the script. The SCRIPT row across the top offers Latin, Chinese, Devanagari, Japanese and Korean. Latin is selected by default; change it before the text comes back if your page is in another script.
- Wait for the extraction. The screen shows a progress state while the recogniser works through the pages. Time depends on page count and phone, not on your connection.
- Fix anything that came out wrong. Tap into the text and edit it. Faint print, stamps and handwriting are where mistakes cluster.
- Copy, share or save it. The three actions in the top bar copy the text to the clipboard, open the share sheet, or write it out as a .txt file.
Questions
Does OCR need an internet connection?
The recognition runs on your phone and neither the page image nor the extracted text is uploaded. One caveat worth stating plainly: the Latin recogniser is delivered as part of Google Play services rather than bundled into the app, so on a device that does not already have that model, Play services fetches it once. The Chinese, Devanagari, Japanese and Korean models are bundled in the app itself. After that first fetch, OCR works in airplane mode.
Which languages can it read?
Five scripts, not five languages — a script covers every language written in it. Latin handles English and the European languages; Devanagari handles Hindi, Marathi and Sanskrit; and there are dedicated recognisers for Chinese, Japanese and Korean. Scripts outside that set, such as Arabic, Tamil or Bengali, are not supported.
Is OCR free?
No. OCR is a Pro tool in the current release. Scanning, signatures, annotation and image compression are among the tools that stay free.
Can I correct the text before I use it?
Yes. The result is an editable field, and copy, share and save all act on the edited version.