F FileShrink

Convert HTML to TXT — strip tags and keep just the content

Drop an HTML file below and get a clean plain-text file back. The browser’s own DOM parser extracts only the readable content. Everything stays on your device.

Your file is processed entirely in your browser. Nothing is uploaded.

HTML files are great for display but terrible for text processing. If you want to feed an article to a language model, run a word count on a web page, index an archive with grep, or copy content into a plain-text editor, HTML tags get in the way. HTML-to-TXT conversion strips the markup and returns just the readable content — the words a reader would actually see if they opened the page in a browser — as a pure .txt file ready for any text tool.

This converter uses the browser’s built-in DOMParser to parse your HTML, then cloning the body into a hidden element and reading its innerText. innerText is smarter than just removing tags: it respects block-level boundaries by inserting line breaks, skips hidden elements, and decodes HTML entities automatically. The result is the same text you would see by selecting everything and copying it out of a browser window. No uploads, no server round-trip, no external parsing service.

How to convert HTML to TXT in 4 steps

  1. 1

    Select your HTML file

    Drop the .html or .htm file into the dropzone or click to browse. Any standard HTML document is accepted.

  2. 2

    Let the DOMParser extract the content

    The tool parses the HTML into a DOM, isolates the body, and reads the rendered innerText — the same text a reader would see on the page.

  3. 3

    Check the result

    The result panel shows the size of the plain-text output. It should be a tiny fraction of the original HTML because all the markup is gone.

  4. 4

    Download the TXT file

    Save the .txt to your device. Feed it to a script, copy it into an editor, or archive it in a grep-able folder.

HTML vs TXT: rendered document versus raw content

HTML is a rendered document format — its files contain markup that tells a browser how to display structure, links, images, styles, and interactivity. A plain text file is just the readable content with no structure beyond line breaks. Converting HTML to TXT is the operation of throwing away everything the browser would render visually and keeping only the words themselves, in roughly the same reading order. It is a lossy conversion in the sense that you lose links, images, headings (as styled elements), and everything that depends on visual rendering. But for text-processing workflows — analysis, search, feeding another tool — that loss is exactly what you want.

When should you convert HTML to TXT?

Feeding a web article to a language model

Most language models want plain text, not HTML with tags mixed in. Converting first gives you clean input that the model handles correctly without tokenizing tag fragments.

Running text analysis on a web page

Word counts, sentiment analysis, keyword extraction, and readability scores all work on plain text. Strip the HTML first and the analysis runs cleanly on the actual content.

Indexing a folder of HTML archives with grep

grep and ripgrep work on plain text. Converting an HTML archive to .txt gives you a searchable corpus without modifying the originals.

Copying article content into a plain-text editor

Sometimes you just want the text without the HTML getting in the way. A .txt version is ready to paste into any text editor, script, or note-taking app.

Frequently asked questions

Why use innerText instead of textContent?
textContent returns every character of every text node without any structure — including content inside <script> and <style> tags, which is usually noise. innerText simulates what a user would copy if they selected everything in the rendered page: it respects block-level boundaries (inserting line breaks between paragraphs), skips hidden elements, and decodes HTML entities. The result reads much better.
Are images and links handled?
Images are dropped entirely — there is no way to represent them in plain text. Links lose their <a> tags but the anchor text is preserved. If you need the link URLs, use the HTML-to-Markdown conversion instead (when we ship it).
What encoding does the output use?
UTF-8 with no BOM. HTML entities like &amp;, &lt;, and &#233; are decoded to their actual characters. Non-ASCII content renders correctly in any modern editor.
Does it handle HTML fragments or only full documents?
Both. Even if your input is just a fragment (a <div> without <html> or <body> wrappers), the DOMParser handles it correctly and extracts the text.
Is my HTML uploaded?
No. The conversion runs entirely in your browser using the built-in DOMParser and innerText APIs. No libraries are loaded for this tool, no files are uploaded, no server is involved.

Related conversions