UTF-8 decoding is the process of converting a sequence of bytes back into the original Unicode characters they represent, and a strict decoder rejects malformed byte sequences with an explicit error instead of replacing them with the U+FFFD replacement glyph. UTF-8 is a variable-width encoding defined by RFC 3629, so each Unicode scalar value uses one to four bytes, with specific leading and continuation byte patterns. When you have a string of bytes shown as hexadecimal, decimal, or 8-bit binary tokens, you can decode them back to text only if every byte fits its role in the pattern. The challenge is that many byte sequences are visually plausible but mathematically invalid, including overlong encodings, truncated multi-byte sequences, isolated continuation bytes, surrogate code points, and values above U+10FFFF. A safe decode workflow treats any of these as a hard failure so you do not present guessed text as verified data. This is exactly the kind of work the UTF-8 Encoder / Decoder handles locally in your browser, using the platform's fatal TextDecoder so bad bytes surface as errors rather than as replacement characters.

Why UTF-8 Decoding Breaks in the Real World
UTF-8 looks simple at first glance, but the byte patterns are unforgiving. A valid UTF-8 stream is a sequence of bytes where every leading byte announces how many continuation bytes follow it, and every continuation byte starts with the bits 10. When any byte breaks that contract, the stream is no longer well-formed UTF-8 even if the bytes happen to map to printable characters in another encoding.
Several specific failure modes appear over and over in log files, network captures, and exported data:
- Truncated sequences at chunk boundaries, such as the bytes E2 82 with no third byte when a transfer was split.
- Isolated continuation bytes left over when a leading byte was lost or transformed.
- Overlong encodings that use more bytes than necessary, like C0 AF for the slash character.
- Surrogate code points in the U+D800 to U+DFFF range, which are reserved for UTF-16 and must never appear in UTF-8.
- Scalars above U+10FFFF, which exceed the defined Unicode range entirely.
A non-strict decoder substitutes U+FFFD for every one of these failures, leaving a string of replacement characters in the output. That hides the original problem because the user sees text that almost works and may not realize the bytes were corrupted. A strict UTF-8 decode instead returns an explicit error the moment it sees an invalid sequence, which is the only way to be certain that what you are reading actually came from the original source rather than from a guess.
The Three Byte Notations You Can Decode From
Before you can decode, you need to know which representation your source data uses. UTF-8 itself is a single byte sequence, but humans almost never read raw bytes, so the bytes get written in one of three notations:
- Hexadecimal uses uppercase two-digit bytes separated by spaces (for example 24 41 C2 A2). It also accepts space- or comma-separated one- or two-digit byte tokens with optional 0x prefixes, or one continuous even-length hexadecimal string such as 2441C2A2.
- Decimal uses integer tokens from 0 through 255 separated by spaces (for example 36 65 194 162). It accepts integer tokens only and does not interpret letters.
- Binary uses exactly eight zero-or-one characters per token (for example 00100100 01000001 11000010 10100010).
These are three views of the same byte array. Switching notation does not change the underlying UTF-8 sequence, so a correct decode produces the same text regardless of which representation you started with. Picking the right one matters only for the parser: hex tolerates continuous strings and 0x prefixes, decimal accepts integer tokens, and binary is strict about exactly eight bits per token. If you paste bytes in a format the parser does not recognize, the decode fails at the token stage before the byte sequence is even examined.
How to Decode UTF-8 Bytes Step by Step
- Open the UTF-8 Encoder / Decoder in your browser tab.
- Choose the UTF-8 bytes to text direction so the tool knows you are decoding, not encoding.
- Select the notation that matches your input: hexadecimal, decimal, or binary. If you are unsure, start with hexadecimal because it tolerates continuous strings, spaced tokens, comma separators, and optional 0x prefixes.
- Paste the byte sequence into the input area. The parser accepts space- or comma-separated tokens with optional 0x prefixes, and one continuous even-length hexadecimal string.
- Select the conversion button to run the decode. Conversion runs entirely in the current tab, so nothing is uploaded or stored.
- Compare the decoded text against the original source format. If the source claimed to be UTF-8 and the output matches, the round trip is good.
- To confirm the result is stable, encode the decoded text back to bytes and check that you get the exact same byte sequence you started with before you overwrite the original data.
The tool is bounded at 200,000 bytes of decoded notation and 200,000 UTF-16 code units of text, so keep your input within those limits. If you have a much larger file, use a dedicated binary tool that streams from disk rather than loading everything into memory.
Reference Byte Patterns for Common Characters
Every valid UTF-8 character follows one of four byte-length patterns, and the leading byte always announces which pattern is in use. The reference table below uses verified values from RFC 3629 and the Unicode Standard core specification.
| Code point | Character | Byte length | UTF-8 bytes (hex) |
|---|---|---|---|
| U+0024 | $ (dollar sign) | 1 | 24 |
| U+0041 | A (Latin A) | 1 | 41 |
| U+007F | DEL (boundary) | 1 | 7F |
| U+00A2 | ¢ (cent sign) | 2 | C2 A2 |
| U+0800 | Samaritan letter boundary | 3 | E0 A0 80 |
| U+20AC | € (euro sign) | 3 | E2 82 AC |
| U+1F600 | 😀 (grinning face) | 4 | F0 9F 98 80 |
| U+10FFFF | Maximum scalar | 4 | F4 8F BF BF |
The dollar sign and the Latin letter A sit inside the 1-byte ASCII range, the cent sign crosses the U+0080 boundary into 2-byte territory, the euro sign uses 3 bytes, and the grinning face emoji uses all 4 bytes. The maximum scalar U+10FFFF encodes to F4 8F BF BF, which is the largest well-formed UTF-8 sequence that can ever exist under the standard.
What Happens When a Decode Fails
A strict UTF-8 decoder fails loudly on bad input instead of guessing. Each row below describes an input shape and the failure you should expect when you paste it into a fatal decoder.
| Input shape | Example bytes | Failure mode |
|---|---|---|
| Overlong encoding | C0 AF | Rejected: slash does not need 2 bytes. |
| Truncated sequence | E2 82 | Rejected: leading byte promises 3 bytes but only 2 appear. |
| Isolated continuation | 80 alone | Rejected: a continuation byte has no leading byte. |
| Surrogate encoding | ED A0 80 | Rejected: U+D800 to U+DFFF are not valid UTF-8 scalars. |
| Out of range | F5 80 80 80 | Rejected: encodes a value above U+10FFFF. |
| Empty input | (nothing) | Decodes to empty text without an error. |
The same rule applies when you encode in the other direction. If the input text contains an unpaired UTF-16 surrogate code unit, the tool rejects it before calling TextEncoder so a claimed lossless conversion cannot silently change the original input. Valid surrogate pairs that represent real supplementary characters, such as emoji, are accepted and encoded correctly into four UTF-8 bytes.
Verify the Round Trip Before Trusting the Result
A single decode pass tells you what the bytes should say, but it does not prove that you have the right bytes in the first place. The standard discipline is a round-trip check: decode the bytes to text, then encode the same text back to bytes and compare the new byte sequence against the original. If the two byte sequences match, the decode is lossless and the representation is consistent. If they differ, you either had malformed bytes to begin with, or you picked the wrong notation for the parser.
Run the round trip on a small, known sample first. A handful of representative characters, including one ASCII character, one 2-byte character, one 3-byte character, and one 4-byte emoji, is enough to expose most bugs in the byte boundaries. Only after the round trip succeeds should you replace the original source. The round-trip verification guide for UTF-8 browser tools goes deeper into this discipline, including how to keep the original data intact until you have a verified output.
If the bytes came from a system that might have used Windows-1252, Shift JIS, GBK, or any ISO-8859 family, do not force them through UTF-8. The page assumes UTF-8 only and rejects bytes that look valid but decode to gibberish, which is a clue to identify the original encoding before continuing rather than guessing.
For a deeper look, see Vigenere Cipher Decoder Bulk: Decode Long Ciphertexts.