PDF to Text
Pull the text layer out of a PDF. If a document returns nothing, it is a scan — the explanation below covers what to do about that.
How to use the PDF to Text
- Drop in the PDF you want to read.
- Choose which pages, and whether to keep the original line breaks or reflow into paragraphs.
- Press Extract text, then copy the result or download it as a text file.
Why some PDFs return no text at all
There are two completely different kinds of PDF that look identical on screen.
A digital PDF — exported from Word, a browser or a design tool — stores text as characters with positions and fonts. Extraction reads those directly and is essentially perfect.
A scanned PDF is a photograph of paper. There are no characters in the file, only pixels arranged to look like letters. No extraction tool can read it, because there is nothing to read. If this tool reports zero words, that is what you have.
The way forward is OCR — optical character recognition — which analyses the image and guesses at the letters. Many PDF readers include it, and some scanners apply it automatically, producing a "searchable PDF" that has both the image and a hidden text layer. Those extract fine here.
Why extracted text sometimes reads oddly
PDF is a layout format, not a document format. It records where each piece of text sits on the page, not which paragraph or column it belongs to. Reconstructing reading order is guesswork, and certain layouts defeat it:
- Multiple columns can interleave, producing lines that alternate between columns.
- Tables lose their structure — cells come out as a run of values with no rows.
- Headers, footers and page numbers appear mixed into the body text.
- Hyphenated words split across lines may stay broken.
- Ligatures like fi and fl sometimes extract as unexpected characters.
The "flowing paragraphs" option rejoins lines that appear to continue a sentence and repairs hyphenation. For anything with columns or tables, "preserve line breaks" gives you more to work with.
A note on copying from PDFs
Text extraction gives you the words, not the rights to them. A PDF may be copyrighted, licensed, or confidential regardless of how easy the text is to copy. Some documents also carry technical restrictions on copying, which this tool does not attempt to bypass — encrypted files simply will not open.
Frequently asked questions
Why did I get no text from my PDF?
It is a scan — an image of pages rather than a text document. There is genuinely no text in the file to extract. You need OCR, which is available in Adobe Acrobat, many scanner apps and several free tools.
Does it do OCR?
No. OCR needs a trained recognition model, which is a substantial download and a different kind of processing. This tool reads the text layer that is already in the file.
Can it preserve formatting?
Not bold, italics or layout — the output is plain text. Line breaks can be preserved, which helps with poetry, code and structured documents.
Why are the columns jumbled?
Reading order is inferred from position on the page, and multi-column layouts are ambiguous. Try "preserve line breaks", or extract one column at a time by cropping the PDF first.
Is my document uploaded?
No. Extraction runs in your browser using pdf.js, so confidential material never leaves your device.
Related tools
Convert PDF pages into JPG or PNG images.
Pull selected pages out of a PDF, or delete pages.
Count words, characters, sentences and reading time as you type.
Split a PDF into separate files by range or page count.
Sort lines alphabetically, numerically, by length or at random.