F FileShrink

Convert PDF to TXT — extract all text from your PDF

Drop a PDF below and get a clean .txt file with the text content from every page. The extraction runs in your browser via pdfjs — no OCR, no upload, no server.

Your file is processed entirely in your browser. Nothing is uploaded.

PDFs are designed for visual fidelity, not for text processing. When you need to feed a PDF’s content to a script, run a word count, paste it into a plain-text editor, search across documents with grep, or use it as input for a language model, you first need to get the text out of the PDF container. For PDFs that were created from digital sources (Word, LaTeX, InDesign, most web-to-PDF tools), the text is embedded as selectable characters, and extracting it is straightforward. This tool does exactly that: it reads every page of your PDF with pdfjs, pulls out the text items, and concatenates them into a single .txt file.

The extraction is not OCR — it reads the text objects already stored in the PDF, which is fast and accurate for digitally-created documents. If your PDF was produced from a scanner or is a photograph of text, the embedded text layer may be empty or missing, and this tool will return little or no text. For scanned PDFs, you would need an OCR tool (which we do not offer yet). The full extraction runs in your browser through pdfjs-dist, so your PDF never leaves your device and there is no server involvement.

How to extract text from a PDF in 4 steps

  1. 1

    Select your PDF file

    Drop the .pdf into the dropzone or click to browse. Any PDF with embedded text is accepted — there is no file size limit from our side.

  2. 2

    Let pdfjs read every page

    The library opens your PDF, iterates over each page, and calls getTextContent() to extract the text items that the PDF renderer would normally display.

  3. 3

    Wait for text concatenation

    Text items from each page are joined into paragraphs, and pages are separated by blank lines. A typical 10-page document finishes in under a second.

  4. 4

    Download the TXT file

    Save the .txt to your device. The text is UTF-8 encoded and ready for any text tool, script, editor, or language model input.

PDF vs TXT: layout-perfect versus text-only

A PDF preserves the exact visual layout of a document — every character is positioned at specific coordinates, fonts are embedded, and the file looks identical on every viewer. A TXT file strips all of that away and keeps only the characters themselves, in reading order, with line breaks between logical sections. Converting PDF to TXT is the operation of throwing away the visual container and keeping only the words. You lose formatting, fonts, images, tables (as visual elements), headers and footers, and page numbers — but you gain a file that any text tool on the planet can process. For content analysis, search indexing, and pipeline inputs, this trade is almost always worth it.

When should you convert PDF to TXT?

Feeding a PDF report to a language model

LLMs want plain text. Extracting the text from a PDF first gives you a clean input without PDF structure artifacts mixed in.

Running word count or content analysis

Word counters, keyword extractors, and readability tools all need plain text. Extracting from PDF first ensures you are analyzing real content, not markup.

Indexing a PDF archive with grep or full-text search

Converting a folder of PDFs to .txt makes the whole archive searchable with standard command-line tools. grep does not read PDFs, but it reads .txt perfectly.

Copying the content into another editor or format

Sometimes you need the words from a PDF in a Google Doc, a markdown file, or a CMS editor. Extracting to TXT first gives you clean copyable content without phantom formatting.

Frequently asked questions

Does this use OCR?
No. This tool reads the text objects already embedded in the PDF file structure. If the PDF was created digitally (from Word, LaTeX, InDesign, a web page, or any print-to-PDF workflow), the text is there and extraction is fast and accurate. If the PDF is a scanned image with no embedded text layer, the output will be empty or near-empty. For scanned PDFs, you need a dedicated OCR tool.
Does it extract text from every page?
Yes. The tool iterates over every page of the PDF and concatenates the text. Pages are separated by blank lines in the output. A 100-page PDF produces a single .txt file with all pages in order.
Is the text order correct?
For most PDFs, yes. pdfjs reads text items in the order the PDF creator placed them, which is usually left-to-right, top-to-bottom reading order. Complex multi-column layouts may produce text in an unexpected order because PDF does not have a concept of columns — it only has positioned glyphs.
What about tables in the PDF?
Tables in a PDF are not semantic structures — they are just text positioned at grid coordinates. The extraction will produce the cell contents as a stream of text in reading order, but without column delimiters or row breaks. For structured table extraction, you would need a specialized PDF table parser.
Is my PDF uploaded?
No. pdfjs-dist runs entirely in your browser. The PDF is parsed locally, text is extracted locally, and the .txt file is produced locally. Nothing is ever uploaded, logged, or stored.

Related conversions