A text to hex example shows each character of a string rewritten as the two-digit hexadecimal representation of its UTF-8 byte, so the ASCII letter "H" becomes the byte 48 and "i" becomes 69, producing the continuous output 4869 for the input "Hi". The mapping comes from the WHATWG Encoding Standard, which defines UTF-8 as a variable-width encoding where one byte carries ASCII, two bytes carry most accented Latin letters, three bytes carry the bulk of the Basic Multilingual Plane, and four bytes carry supplementary characters such as emoji. A text to hex example therefore needs to show not just the letters typed but the byte count that results, because a four-character emoji string occupies far more hex digits than four ASCII letters. The Text To HEX tool applies this exact UTF-8 byte mapping in the browser using the standardized TextEncoder API, so the numbers shown in any text to hex example reflect the bytes a strict decoder would read.

Reading the Output: Byte Count, Replacement Count, and Output Length
A typical text to hex example includes three pieces of supporting data alongside the digits: the UTF-8 byte count, the count of isolated surrogate code units that were replaced with U+FFFD, and the formatted output length. The byte count tells how many two-digit hex pairs the result contains in plain format. The replacement count is non-zero only when input contains an unpaired UTF-16 surrogate code unit, which is rare in typed text but can appear in malformed or programmatically generated strings. The formatted output length depends on the chosen format and is calculated before the output string is built. Plain format writes 2n digits, where n is the byte count. Spaced format writes 3n−1 characters because one ASCII space separates each pair. 0x-prefixed format writes 5n−1 characters because each token is "0xNN" followed by a space except after the final token.
Walkthrough: Encoding a Short String Step by Step
Take the input "Hi". The browser receives two UTF-16 code units, one per character. TextEncoder converts this to two UTF-8 bytes because both characters fall inside the ASCII range, so each occupies a single byte. The byte for "H" is 48 and the byte for "i" is 69 in hexadecimal. In plain format, the result is the continuous string 4869. In spaced format, the result is "48 69" with one ASCII space between the two pairs. In 0x-prefixed format, the result is "0x48 0x69" with no delimiter after the final token. Using the format length formulas with n equal to 2, plain gives 2×2 = 4 characters, spaced gives 3×2−1 = 5 characters, and 0x-prefixed gives 5×2−1 = 9 characters. The byte count displayed is 2 and the formatted output length matches the chosen format. None of the three formats changes the bytes themselves; only the punctuation and spacing change.
How to Convert Text to Hex in Three Steps
- Enter the source text in the input field, including any Unicode characters, whitespace, NUL bytes, or punctuation the field accepts.
- Choose the output format (continuous, space-separated, or 0x-prefixed) and select lowercase or uppercase hex digits.
- Encode the text, review the byte count and any isolated-surrogate replacement warnings, then copy the complete hexadecimal output.
Format and Letter Case at a Glance
Changing the format or letter case never alters the encoded UTF-8 bytes; it only changes the formatted output length and the appearance of the digits. The table below summarizes the three formats using the same input bytes:
| Format | Token form | Length formula | Example for "Hi" |
|---|---|---|---|
| Continuous | Two digits per byte | 2n | 4869 |
| Spaced | Two digits per byte separated by ASCII space | 3n−1 | 48 69 |
| 0x-prefixed | "0xNN" per byte separated by ASCII space | 5n−1 | 0x48 0x69 |
For a larger worked case, the format length formulas can be applied directly to a multi-byte input. Consider a single-byte ASCII character such as "A" (byte 41). Plain format yields 2×1 = 2 characters, spaced format yields 3×1−1 = 2 characters, and 0x-prefixed format yields 5×1−1 = 4 characters. Plain and spaced collapse to the same length for a single byte because there are no internal separators to add. For a four-byte emoji such as U+1F600, plain yields 2×4 = 8 characters, spaced yields 3×4−1 = 11 characters, and 0x-prefixed yields 5×4−1 = 19 characters. The plain and spaced columns differ by exactly n−1 ASCII spaces, and the 0x-prefixed column adds 2n characters over the spaced form. A reader who needs to look up the actual byte values for many different inputs can consult the Text to Hex Cheat Sheet: UTF-8 Values and Format Reference for a broader reference.
Examples by Character Type
Different character classes produce different lengths in any text to hex example. The mapping is defined by the UTF-8 spec rather than by the tool, so the same input always produces the same bytes:
- ASCII such as U+0041 "A" encodes as one byte, 41.
- Accented Latin such as U+00E9 "é" encodes as two bytes, c3 a9, because 0xE9 falls outside the ASCII range but inside the two-byte window.
- CJK within the Basic Multilingual Plane such as U+4F60 "你" encodes as three bytes, e4 bd a0, because the codepoint sits above the two-byte range.
- Supplementary characters such as U+1F600 "😀" encode as four bytes, f0 9f 98 80, because the codepoint exceeds U+FFFF.
Uppercase mode changes only the letters a through f to A through F; the byte values themselves and the lowercase 0x marker remain the same regardless of the case setting. So 0x1f600 written in uppercase is 0x1F600, and the byte 0xc3 is C3 in uppercase mode.
Why Some Bytes Look Unexpected in a Text to Hex Example
Several inputs produce hex that surprises readers who expect simple one-byte-per-character behavior. The NUL character U+0000 encodes as the single byte 00, not as nothing. CR, LF, tab, and space encode in whatever order they appear in the input and are not normalized; CR LF stays as 0d 0a rather than collapsing to 0a. Combining marks are not folded into their composed form, so the sequence "e" followed by U+0301 (combining acute accent) encodes as 65 cc 81 rather than being changed to the precomposed "é" U+00E9. An isolated UTF-16 surrogate is not a valid Unicode scalar, and the standardized encoder replaces each occurrence with U+FFFD, whose UTF-8 bytes are ef bf bd, while the tool surfaces a visible replacement count because the bytes cannot decode back to the original isolated code unit. Finally, the encoder does not prepend a BOM; if your input itself begins with U+FEFF, those three bytes ef bb bf appear as ordinary data because the character is part of the input string. The browser API used for the conversion is described in the MDN TextEncoder.encode reference.
Round-Tripping a Text to Hex Example Through the Companion Decoder
A clean round trip requires the bytes to decode back to the original string. The companion Hex to Text tool reads those bytes with explicit separator and 0x-prefix rules and decodes only well-formed UTF-8. Because TextEncoder does not add a BOM, an input that begins with U+FEFF produces leading bytes ef bb bf, which the companion decoder consumes under its documented BOM behavior rather than treating as ordinary data. For exact round trips, the original text should not start with U+FEFF and should contain only well-formed Unicode without isolated surrogates. Other data passes through unchanged: the tool does not apply normalization, newline conversion, trimming, case folding, escape parsing, or locale-aware rewriting, so the byte sequence a text to hex example produces is the same one a strict UTF-8 decoder will read.
For a deeper look, see Decode UTF-8 Bytes Without Silent Replacement Characters.