Overview
OCR is the bridge between a picture of a document and a document a computer can read. Everything else — search, extraction, AI question answering — depends on it.
Why a scanned PDF has no text
A scanner does not produce text. It produces an image of a page, and the PDF format stores that image. To a computer, the result is a picture: readable by a person, invisible to search.
What OCR does
- Analyses the page image.
- Recognises characters and words.
- Reconstructs a text layer alongside the original image.
- Makes the content searchable, indexable and reusable.
From scanned pages to answers
- Scanned PDF opened on the device.
- Page images analysed by OCR.
- Text extracted from the pages.
- Document index built from the text.
- AI document chat able to retrieve and answer from the content.
Text extraction
The recognised text can also be exported. PDF JustFit can extract text from a PDF and write it out as a TXT file, which is useful when the content needs to be reused somewhere else.
Accuracy depends on the input
OCR quality depends on the quality of the page images. Clean, straight, high-resolution scans recognise better than skewed or blurred ones.
Key facts
- Automatic OCR for scanned pages
- Runs on-device
- Feeds the document index used by AI chat
- Supports PDF to TXT extraction
Where OCR is used
OCR is what makes a scanned document searchable and what allows it to take part in AI document chat. See PDF OCR for the feature and AI PDF Chat for what it enables.
Frequently asked questions
What does OCR actually do?
OCR analyses the page images in a scanned document and reconstructs a text layer, so the content becomes searchable, indexable and reusable.
Does PDF JustFit OCR automatically?
Yes. Scanned documents are OCR-recognised automatically, so nothing is left out of the document index.
Sources
Last verified: