06 / PDF to text

Extract text from a PDF in the browser.

Pull the real text out of a digital PDF and copy it or save a .txt file. It works for invoices, exports, and documents made in Word. A photo of a page has no text layer, so a scan will come out empty.

How to use PDF to text

  1. 1. Add a digital PDF

    Choose a PDF that you can already select text in, such as an invoice export or a Word-made file.

  2. 2. Read the extract

    Text appears on the page, with a marker for each PDF page. Nothing is uploaded while it is read.

  3. 3. Copy or download

    Copy from the box or download a .txt file. If the box is almost empty, the PDF is a scan.

Digital PDFs and scans are different

A PDF from a billing system or from Word contains characters. This tool reads those characters. A scan, a phone photo, and many 'print to PDF' files from a copier are pictures. They look like text and still return a blank extract. Turning a scan into text needs OCR, which this page does not do.

What the extract looks like

Lines follow the PDF's text layer, which is often reading order rather than the visual columns. Tables may arrive as runs of words instead of cells. It is the right tool for grabbing an address, an invoice number, or a paragraph. It is the wrong tool for rebuilding a spreadsheet.

A private place to copy a paragraph

Pasting a contract into a random website is a bad trade for convenience. Here the file is parsed locally, up to 80 pages, and you can close the tab when the paragraph is copied.

Questions

Why is the extracted text empty?

The PDF is probably a scan or a set of images. This tool reads a text layer only. If you cannot highlight a sentence in a normal PDF reader, there is nothing here to extract.

Can I copy the text without downloading a file?

Yes. The extract is shown on the page. Download the .txt file only if you want to keep it.

Does PDF to text upload my file?

No. Extraction runs in your browser. The PDF is not sent to Matoshri Infotech.

Will columns and tables stay intact?

Often no. You get the words in the order the PDF stored them. Use it to copy text, not to rebuild a table.

Related tools