VitaliWeb Tools

Copy PDF text with readable paragraphs

Inspect source text and page positions, correct the reading order and review cleanup before copying or exporting TXT and Markdown.

Processed on your deviceFree · No account needed

Preparing your tool…

How to use it

  1. Choose or drop a PDF, or load the synthetic study notes.
  2. Suggest reading order and check the source preview. Correct individual line positions and mark headings, lists or protected code/table lines.
  3. Choose paragraph joining and optional repeated-margin removal. Confirm each hard-hyphen break separately; real compound words remain unchanged by default.
  4. Create the cleaned preview, compare the original and change log, and undo up to five recent edits if needed.
  5. Copy the reviewed text or prepare TXT/Markdown. Generated bytes are reopened as UTF-8 and compared before download.

The three-page synthetic PDF contains German columns with Datenver- / arbeitung, a real E-Mail-Konto compound, English paragraphs, lists, code and repeating margins. Confirm only the Datenver- break; code and real hyphens must remain.

Details & limits

Selectable text only; no OCR. Up to 25 MiB, 100 pages, 50,000 fragments and 2 million text characters. A 60-second worker limit applies to each operation. Complex compression and unusually large expanded contents may be rejected. Text export is limited to 16 MiB. Columns, paragraphs and margin repetition are heuristic; tables, code and rotated text need review. Embedded images: unfiltered/Flate pixels or checked JPEG, up to 8,192 px per side. Inline images, JPEG 2000, JBIG2, CCITT and unusual JPEG variants are unsupported.

Will it correctly reconstruct every column, table or scanned page?

No. Reading order and cleanup are inspectable suggestions. Reorder lines and protect complex content manually. Pages without extractable text are listed separately; scans and outlined letters are not claimed as successful extraction.