Copying text out of a PDF by hand is miserable: line breaks in odd places, columns jumbled together, and one page at a time. This tool pulls the text out of the whole document in one go and puts it in a box you can edit, copy or save as a .txt file. It reads the file in your browser, so the contents never leave your device.
What it's good for
- Reusing paragraphs from a report or brochure without retyping.
- Pasting a long document into a translation tool or a proofreading app.
- Searching a big PDF for a word in a simple text editor.
- Getting the wording of a contract into an email so you can discuss a clause.
- Checking the text before you sign; it's much easier to read in a plain box than in a scanned layout.
Using the result
Each page's text is separated by a marker such as "--- Page 3 ---", which you can switch off if you want continuous text. Copy everything with one button, or download a .txt file named after the PDF. Paste it into a document, then tidy up the formatting. Some PDFs use hard line breaks at the end of every line, so you may need to join lines; a find-and-replace for line breaks inside paragraphs usually does the trick.
Scanned PDFs are different
If the PDF was made by scanning paper or photographing pages, there's no actual text in it, only pictures of text. In that case this tool will find little or nothing, and you'll need OCR (optical character recognition) software. A quick test: if you can't highlight a word in your PDF reader, it's a scan. The tool tells you when it thinks that's the case instead of leaving you with an empty box. If you only need an image of a page, PDF to JPG is the answer.
Layout, columns and tables
Plain text can't hold columns, tables, fonts or images, so the result is the words in reading order with a separator between pages. It's ideal for copying and editing, not for recreating the original design. Multi-column layouts sometimes come out one column after the other, which is usually what you want, and tables turn into rows of values separated by spaces. If a document has complicated formatting, extract the text, then format it again in your word processor.
What to do with the text
Once you have the words, there are plenty of ways to use them. Check the length with the word counter, fix capitalization with the case converter, or turn the edited text into a fresh document with Text to PDF. For documents that you need to number or mark, Add Page Numbers and Watermark PDF finish the job.
Privacy and accuracy
Everything happens in your browser, so confidential contracts and personal letters stay on your device. Accuracy depends on how the PDF was made: files exported from Word or a web page give almost perfect text, while PDFs with unusual fonts can produce odd characters. Check any important figures against the original, particularly numbers, dates and names. If the text contains details you should not share, use Redact PDF on the original before you send the file anywhere.
Password-protected files need to be unlocked in your reader first. Very long documents may take a few seconds, and the page counter shows progress as it goes.
A worked example
A colleague emails you a 30-page policy document and asks you to summarize clause 7. Open it here, wait a moment while the pages are read, and press Copy all. You can paste the whole text into an email or a note-taking app, search for "clause 7", and copy just the passage you need. Without this step you'd have to scroll through the PDF, drag to select text across pages and fix broken line endings by hand. With the page markers on, you also know which page each paragraph came from, which is useful when you cite it.
Cleaning up extracted text
Text copied from a PDF often needs a little tidying. Hyphenated words at line ends, such as "inter- national", can be joined with a find-and-replace for the hyphen and space. Page headers and footers repeat on every page, so delete them using a find on the repeated phrase. Bullet symbols may turn into odd characters; replace them with dashes. Spend a minute on cleanup and the text is ready for editing, translating or reading aloud.
Troubleshooting
If words are missing, the PDF may mix text with scanned images, and the images won't be read. If the text comes out in a strange order, the PDF's reading order is unusual, so check it against the original. If letters look wrong, the file may use a custom font whose characters aren't mapped properly; that's a limitation of the PDF itself. For important text, always compare against the original before you rely on it.