PDF to Text: what it does and how to use it
Table of contents
Extracting text from a PDF gives you the words without the formatting - no layout, no images, no styling, just content you can paste, search or feed into another tool. Whether it works at all depends on one thing: whether the PDF contains real text or pictures of text. A PDF exported from Word, InDesign or a web page stores actual characters, and extraction is instant and perfect. A PDF built from scanned or photographed pages stores images, and there are no characters to extract - you would get nothing back. The quick test takes two seconds: open the PDF and try to select a line with your cursor. If individual words highlight, extraction will work. If you can only draw a box over the page, it is a scan and needs OCR instead. The second thing to expect is reading order. A PDF records where each character sits, not that this is column one and that is column two, so multi-column layouts often extract in an order that reads oddly. This runs in your browser.
PDF tools that upload to a server are a poor fit for contracts, IDs and anything you would not paste into a random website. This one is meant for that exact case.
PDF to Text is a good fit when copying content out of a PDF that will not let you select text.
The useful part
PDF to Text is built around a few practical wins, not a long feature list:
- Instant and lossless for any PDF with a real text layer - no recognition errors, because nothing is being recognised.
- Gets around PDFs that block selection or copying in a viewer.
- Runs in your browser, so contracts and reports never leave your device.
- Output is plain text, which every other tool accepts.
- No account and no page limit.
Do this, in order
- Check the PDF has real text. Try selecting a line in any PDF viewer. If words highlight individually, extraction will work.
- Add your PDF. Drop it in or browse. The file is read locally and never uploaded.
- Extract. All readable text is pulled out as plain text, in the reading order the document stores.
- Copy or save the result. Paste it wherever you need it, or keep it as a plain text file.
Who it is for
- Copying content out of a PDF that will not let you select text.
- Feeding a document into a summariser, translator or word counter that needs plain text.
- Extracting quotes or references from a report or paper.
- Getting the raw text of a contract for comparison against another version.
- Pulling content out of an e-book or manual for searching.
If you want a clean result
- Test selectability first. It takes two seconds and tells you immediately whether this tool or OCR is the right one.
- Expect multi-column documents to extract in an odd order - the PDF stores positions, not column structure.
- If you need just a few pages, split the PDF first rather than extracting everything and hunting through it.
- For a scanned document, use OCR Image to Text instead. It recognises characters from pixels, which is a different operation entirely.
- If you need the layout preserved rather than the words, PDF to Word reconstructs structure instead of flattening it.
Common mix-ups
- Running a scanned PDF through text extraction and concluding the tool is broken. There is genuinely no text in the file to extract.
- Expecting tables to survive. A table becomes a run of words, since the plain text format has no concept of cells.
- Assuming the reading order will match the visual layout on complex pages.
- Using text extraction when you actually wanted an editable document - that is PDF to Word.
- Extracting a whole book when a page range would have done.
Private by default
PDF to Text runs in your browser. The file or text you paste stays on your device. There is no account, and nothing is stored on a ToolBox server for this job.
Related tools worth opening next
If this is one step in a longer job, these usually come after it:
- PDF to Word - Convert PDF to editable Word with accurate layout, headings and tables
- PDF to OCR - Convert scanned PDF to OCR text and a searchable PDF with selectable words
- Word Counter - Count words, characters and sentences
Before you ask
Will this work on a scanned PDF (a PDF made of photographed pages)?
Only if the PDF already has a text layer. Purely scanned/image-based PDFs with no embedded text will return little or no text - you would need an OCR (optical character recognition) tool for those instead.
Does it preserve formatting like bold, tables, or columns?
No - the output is plain text only. Tables and multi-column layouts are flattened into a simple top-to-bottom reading order, which can reorder columns in complex layouts.
Is my PDF uploaded to a server?
No. Extraction runs entirely in your browser, so the document never leaves your device - which matters given how often the PDFs people extract from are contracts, reports and financial statements.
Can I extract text from just some pages?
This extracts the full document. Use Split PDF or the Page Extractor first if you only want specific pages, which is quicker than extracting everything and searching through it.
How do I know if my PDF has real text or is a scan?
Open it and try to select a line with your cursor. If words highlight individually, it has a text layer and extraction will work perfectly. If you can only draw a selection box over the page, it is images and needs OCR instead.
Why is the extracted text in a strange order?
Because a PDF stores where each character sits on the page, not that this block is column one and that is column two. Multi-column layouts, sidebars and footnotes therefore extract in position order rather than reading order.
Open the PDF to Text when you are ready. It is free, and you do not need an account.