The Remove Text Formatting tool does not parse HTML and does not strip tags; literal HTML source stays literal, so the characters <b>literal</b> in the input appear as the same angle-bracket sequence in the output. The tool deliberately rejects any role as an HTML sanitizer or parser because accepting rich markup would mean inserting it into a contenteditable region, executing script, and exposing the page to event-handler, link, image, and layout risks. Instead, the page renders a single native form control — an ordinary textarea — and asks the browser to deliver whatever plain-text representation the clipboard source supplied. If the source application exposes both rich HTML and text/plain on the clipboard, the textarea keeps only the text/plain side. From that point on, the tool does three small things: it validates the string against a one-million Unicode code-point ceiling, normalizes CRLF or standalone CR line endings to LF, and presents the exact same string back to you for copy or TXT download. There is no DOM extraction, no entity decoding, no Markdown conversion, no whitespace collapsing, no smart-quote replacement, no Unicode normalization, and no encoding detection. If the goal is to read HTML as plain text, this is not the right tool — but it is the right tool when the clipboard's plain text is exactly what the source application formatted and you want it back without altering it.

does the remove text formatting tool parse html and strip tags
Does Remove Text Formatting Parse HTML and Strip Tags?

What the Textarea Actually Receives

A native textarea is the boundary between the user's clipboard and this tool. When you paste, the browser consults the clipboard data the source application supplied. The browser keeps the text/plain representation and ignores everything else, including text/html, image/png, and any custom formats. A copy from a word processor, a Google Doc, an email client, a chat app, a source-code editor, or a plain text file all funnel into the same kind of input: a string of Unicode code points with whatever line breaks the source application chose.

The tool never puts this string into a contenteditable element, never creates a DocumentFragment, never parses with DOMParser, and never assigns to innerHTML. That single architectural decision is what keeps the page honest about its limits. It cannot accidentally render a <script> tag because no tag is ever recognized as a tag in the first place — the closing characters </script> are just six Unicode code points sitting inside a JavaScript string. The same reasoning applies to attribute syntax, entity references, inline event handlers, embedded <style> blocks, and inline SVGs. They all arrive as text, and they all stay as text.

What Happens to Literal HTML in the Input

Because the textarea holds a plain string, the characters <, >, /, and letters forming tag-like sequences have no special meaning. Paste the literal source <div class="note">Hello <em>world</em></div> and the result panel, the clipboard payload, and the downloaded TXT file contain exactly those bytes in the same order, with the same spaces and line breaks. The tool does not interpret <em> as emphasis, does not pull world out of the markup, and does not decode &amp; into &.

That preservation is deliberate, because silently rewriting or removing literal tag characters would mislead anyone who pastes source code, HTML email bodies, CMS template fragments, documentation snippets, or forum posts that quote tags as text. The trade-off is unambiguous: if you want only the visible text, you need a true HTML-to-text step before pasting; if you want the literal source preserved, this tool is the correct destination.

Why the Tool Stays Out of HTML Parsing

Parsing arbitrary HTML locally is harder than it looks. Browsers tolerate malformed markup, broken entities, mixed case, embedded SVG, and inline styles, and a naive regex pass can mangle attribute values that contain >, mishandle CDATA sections, or strip angle brackets from code samples the user wanted to keep. A DOM-based extractor has to choose what counts as content versus chrome, and that decision changes the output in ways the user cannot easily preview.

By staying out of parsing entirely, this tool removes those choices. Every accepted character round-trips, the preview equals the copy, the copy equals the downloaded file, and the file equals the Blob's bytes. The deterministic contract — same input string, same output string, only line endings normalized — is the entire product. The trade-off is that the tool can never give you the prose inside an HTML document, and the contract deliberately advertises that limitation rather than papering over it with heuristic stripping.

How to Strip the Other Styling and Keep the Characters

This is the practical sequence when you want unformatted text from a styled source and you do not want any markup characters removed:

  1. Copy the styled text from its source application with the normal keyboard shortcut, confirming that the source supplies a text/plain clipboard representation alongside any rich format.
  2. Open the Remove Text Formatting tool, click into the native textarea, and paste with the same shortcut.
  3. Read the source-supplied preview in the textarea. If the line breaks, list markers, table separators, or trailing whitespace look wrong, the fix lives in the source application — this tool does not rewrite what the clipboard handed it.
  4. Trigger the plain-text creation step. The page validates the string, rejects empty or malformed-Unicode input, rejects anything over one million code points, and replaces CRLF or standalone CR with LF.
  5. Inspect the result panel. Tabs, repeated spaces, blank lines, punctuation, emoji, accented letters, and other scripts should be untouched; literal <tag> sequences should still read as literal text.
  6. Copy the result through the standard Clipboard writeText path — the success message appears only after the browser resolves the write — or download the UTF-8 plain-text file, which uses the text/plain;charset=utf-8 media type and the filename plain-text.txt.

When You Need a Different Tool Instead

The keyword question — does this parse HTML and strip tags — has the same answer regardless of which styled source you started from: no. But the right tool depends on what you actually want at the end.

Input contains Remove Text Formatting output A true HTML-to-text parser would
Literal <b>literal</b> source Same <b>literal</b> characters Output only literal
&amp; entity reference Same &amp; characters Decode to &
<script>alert(1)</script> Same 25 characters Remove or escape tags
Plain prose from a word processor Identical plain prose Identical plain prose

If the goal is to read HTML as prose, the next step is a real HTML-to-text extractor or a guide to removing special characters safely from a text file. Common reasons to reach for a parser instead of this tool include pulling sentences out of a web page source, stripping tags from an email HTML body before pasting into a support ticket, converting CMS export markup into spreadsheet cells, or cleaning chat messages that arrived as rich HTML. If the goal is the opposite — keeping the literal HTML source intact while shedding the source application's own font, color, heading, and link styling — this tool is exactly right, because the clipboard's text/plain representation by definition carries no styles.

You can also reach for related removers when the job is more specific: deleting repeated words, dropping blank lines, collapsing tabs, or trimming trailing whitespace are all jobs this tool deliberately refuses to do, and each one has a focused alternative in the same category.

If you're weighing options, Remove Accents from Text Online Without Losing Letters covers this in detail.