VitaliWeb Tools

Make scanned PDF searchable

Recognise scanned English or German pages and add a positioned searchable text layer.

Processed on your deviceFree · No account needed

Preparing your tool…

How to use it

  1. Choose a PDF or load the synthetic two-page example: one scanned page and one page with existing selectable text.
  2. Choose English, German or both, the page range and recognition resolution. If a scan is sideways, set its clockwise correction in the recognition settings.
  3. Confirm the displayed local engine and language-file download, then recognise the selected pages. Existing text pages are retained without a second layer.
  4. Compare the recognised words and reading order with the source. Names, amounts, umlauts and columns can be wrong; low-confidence words are counted. Try another language or orientation if needed.
  5. Confirm your review, create the searchable PDF, inspect its actual exported preview and download it. Reset or cancel clears the current operation and generated result.

The synthetic scan includes “Hello Zurich: 42 green folders.” and “Grüsse aus Zürich: Öl, Käse, Äpfel.” German recognition can handle these umlauts; accuracy still depends on the selected model. The second page already contains selectable text and must be skipped, with its original text retained once.

Details & limits

Up to 25 MiB input/output and 100 pages; 60 seconds per operation. Signed PDFs, XFA, embedded files and active document actions are rejected. Checked JPEG and unfiltered/Flate images are supported within decoded-resource limits; inline images, JPEG 2000, JBIG2 and CCITT are unsupported. No PDF/A, PDF/UA or signature-validity certification. At most 10 OCR pages, 6 megapixels per raster, 8,192 pixels per side, 20,000 recognised words and 200,000 characters. Recognition at 150 or 200 dpi, subject to pixel limits. Pages with any existing text are skipped, including mixed text-and-scan pages. The selected engine/model download size is shown before your confirmation. Characters outside the supported Unicode layer are rejected instead of silently omitted.

Does OCR change the scan or guarantee correct text and accessibility?

The visible source pages are retained. OCR adds an invisible Unicode text layer with positioned word boxes; the text can contain recognition or reading-order errors. This workflow does not promise character-perfect highlighting, automatic language detection, OCR of scan regions on pages that already contain text, accessible document tagging or PDF/A conformance. Orientation correction affects recognition only; it does not rotate the exported page.