Scanned PDF or text PDF: choose the right next step

Published by TDS Document Scout · Updated

A PDF can contain actual text, page images or both. Being able to see words does not mean those words are extractable or usable with assistive technology.

Extract a page of text. If the result is empty, inspect the page visually: it may be scanned, blank or contain unsupported content. Empty extraction alone does not prove it is a scan.

For a scan, run OCR in a suitable editor, then check names, numbers, punctuation and reading order. OCR output can contain convincing errors.

A text layer is only the first step. Review headings, alternatives and structure separately. This tool does not perform OCR or certify the resulting PDF.

PDF to text

Reference material

Understand the next step

Useful to someone else?

Share this page with your team. The link opens the tool or guide, with an empty workspace.