Hindi · How-to
Hindi OCR: turn a scanned Hindi PDF into text you can copy
Make a scanned Hindi PDF or photo searchable and copyable, free and on your own device. Steps, tips for better accuracy, and what Hindi OCR can't do.
HarshFounder, DoYourPDF2 min read
To get text out of a scanned Hindi document, run it through OCR (optical character recognition) with Hindi selected. OCR Searchable PDF does this free, on your own device: it reads the Hindi and English on each page and lays the text invisibly over the scan, so you can search, select and copy it.
Steps
- Open OCR Searchable PDF.
- Choose the scanned PDF, or photos of the pages (JPG or PNG).
- Under Language, pick Hindi, or Both if the pages mix Hindi and English.
- Leave Accuracy on Best.
- Press Make searchable and download the result.
The first run downloads the recognition engine (about 4 MB) and the Hindi language data once; after that, your browser keeps them. The scan itself is never uploaded.
Get the text out
The result looks exactly like your scan, with the words now selectable. From there:
- Copy a passage: open the PDF and select the text.
- Get all the text: run the result through PDF to Text.
- Edit it as a document: run the result through PDF to Word.
How to get better results
The engine reads what it can see, so a better scan beats any setting (Tesseract):
| Do | Why |
|---|---|
| Scan or photograph at 300 dpi or more | Small matras and conjuncts need detail |
| Keep the page flat and straight | Skewed lines break words apart |
| Use even light, no shadows | Shadows read as ink |
| Crop away the table edge, fingers and margins | Stray marks become stray letters |
| Pick Both for mixed Hindi and English | Each language needs its own model |
Expect a few mistakes on clear printed pages and many on poor ones. In one study comparing OCR engines on Indian scripts, Tesseract, the engine DoYourPDF uses, read about 95% of Hindi characters correctly, against about 97% for Google's paid cloud OCR (study). Proofread anything that matters, especially numbers and names.
What Hindi OCR can't do
- Handwriting. The engine is trained on printed text.
- Garbled typed text. Some PDFs were typed in older Hindi fonts such as Kruti Dev, which store Hindi letters as English ones, so copied text comes out as nonsense (Digital Orientalist). That is a font problem, not a scan.
- Other Indian scripts. DoYourPDF's OCR reads Hindi and English today.
Questions people ask
Can I convert a scanned Hindi PDF to text for free?
Yes. OCR Searchable PDF recognises Hindi and English on your own device, for free, and adds the text to the PDF. Then copy it, or export it with PDF to Text or PDF to Word.
How accurate is Hindi OCR?
On clear printed pages, quite accurate. In one published study, the Tesseract engine that DoYourPDF uses read about 95% of Hindi characters correctly. Blurry, skewed or handwritten pages do much worse.
Does it work on handwritten Hindi?
Not well. The engine is trained on printed text, so handwriting comes out with many mistakes.
Why does Hindi text copied from some PDFs come out as gibberish?
Those PDFs were typed in older Hindi fonts such as Kruti Dev, which store Hindi letters as English ones. That is a font problem, not a scanning problem, so OCR settings won't fix the copied text directly.
Is my document uploaded?
No. The OCR engine and its Hindi language data download to your browser once, then run on your device. Your scan never leaves it.