Text to hex is the process of writing every character in a string as the two-digit hexadecimal code of its UTF-8 bytes, so the letters "Hi" appear as 48 69 and the emoji 😀 appear as F0 9F 98 80. The conversion always runs in a fixed three-stage pipeline: each Unicode character is identified as a scalar value, that scalar is encoded into one to four UTF-8 bytes using the rules defined in the WHATWG Encoding Standard, and every byte is then written as two hexadecimal digits. Two-digit pairs come from the fact that one byte holds values 0 to 255, which fit into exactly two hex digits; bytes above 0x7F fall back to ASCII letters A through F depending on the chosen letter case. The output preserves byte order and character order, so hex is not encryption or hashing — anyone with the same UTF-8 bytes can rebuild the same text, subject only to the leading-U+FEFF caveat documented for matching decoders. Understanding those three stages is what separates a quick hex dump from a conversion you can trust across accents, ideographs, emoji, NUL bytes, and combining marks.

text to hex explained
Text to Hex Explained: How UTF-8 Bytes Become Hex Pairs

How UTF-8 Text Becomes Hexadecimal

The first stage of text-to-hex is scalar resolution. JavaScript strings are sequences of UTF-16 code units, and well-formed Unicode text groups high and low surrogates into surrogate pairs for characters outside the Basic Multilingual Plane. A valid pair represents one supplementary scalar, while a lone high or low surrogate is not a Unicode scalar at all — the WHATWG TextEncoder standard defines what happens to those lone code units, and the MDN reference documents the matching browser behaviour.

The second stage is byte encoding. UTF-8 is a variable-width encoding: one byte for code points 0 to 127, two bytes for 128 to 2047, three bytes for 2048 to 65535, and four bytes for 65536 to 1114111. The mapping is deterministic and reversible for any well-formed scalar, which is why hex is a lossless view of the underlying string once the input is well-formed.

The third stage is hex formatting. Every byte is rewritten as two characters drawn from 0-9 and a-f (or A-F in uppercase mode). There is no compression, no escape parsing, no locale rewriting, and no trimming — formatting is a presentation step that runs after the UTF-8 byte array is already complete. Encoding, sizing, formatting, display, and clipboard preparation all happen locally in the current browser tab, so nothing is uploaded to round out the process.

How Many Bytes Each Character Takes

UTF-8 byte width depends entirely on the code point, not on the glyph. That is why a four-character sentence can encode into eight, fifteen, or more hex digits without any obvious pattern at first glance. The table below shows one reference scalar from each of the four width bands so the rule is concrete:

Code pointCharacterBytesUTF-8 hex (uppercase)
U+0041A141
U+00E9é2C3 A9
U+4F60你3E4 BD A0
U+1F600😀4F0 9F 98 80

The pairing always reads high nibble then low nibble, which is why U+0041 is "41" rather than "14", and why U+00E9 spans two bytes as "C3 A9" rather than a single token. With this byte-to-pair mapping fixed, what changes between tools is not the encoded bytes but only the separators and letter case around them.

How to Convert Text to Hex with Text To HEX

The Text To HEX tool runs the same three-stage pipeline above directly in your browser using the standard TextEncoder API, so you can move from a typed string to copy-pasteable hex pairs without uploading anything. Use it whenever you need an exact UTF-8 hex view for debugging, protocol work, or sharing bytes with another tool. Follow these steps to produce a reliable result.

  1. Enter your text in the input field, including any Unicode, whitespace, or NUL data the browser field can hold — the encoder works against whatever string the field contains at encode time.
  2. Pick one of three output formats: continuous pairs such as 4869, space-separated pairs such as 48 69, or 0x-prefixed tokens such as 0x48 0x69. Optionally choose lowercase or uppercase A-F.
  3. Click encode. The tool reports the exact UTF-8 byte count, the formatted output length, and — when applicable — the number of isolated surrogates that were replaced with U+FFFD.
  4. Review the replacement warning if any code units were dropped. Without that warning, your hex describes the original text losslessly; with it, you cannot recover the original isolated surrogate code units on decode.
  5. Copy the complete hexadecimal output to your clipboard with the copy button. If clipboard permission is denied, select the read-only result manually; the bytes themselves are unaffected.

Editing the input, changing format or case, or starting a new encode clears the previous result and any stale copy status, so the values you copy always match the current encode.

Reading the Three Output Formats

The three formats are presentation choices over the same byte array. Plain format concatenates two digits per byte and produces a string of length 2n where n is the UTF-8 byte count. Space-separated format inserts one ASCII space between every pair, yielding length 3n minus 1 because no delimiter is added before the first byte or after the last. The 0x-prefixed format writes each byte as "0x" plus two digits and separates tokens with a single space, giving length 5n minus 1.

These exact formulas matter for the output budget. Text To HEX accepts at most 4,999,999 UTF-16 code units in the formatted result, and the boundary is intentionally tight: one million ASCII input characters in 0x-prefixed format produce exactly 4,999,999 output code units. Multi-byte Unicode in a verbose format can hit that ceiling before the input budget is exhausted, in which case the tool rejects the request rather than switching format, slicing bytes, or sampling content. Lowercase and uppercase A-F produce the same byte count and the same separator count, so letter case does not change formatted output length at all.

What the Encoder Preserves and What It Leaves Alone

Text To HEX is intentionally a thin view over UTF-8. NUL (U+0000) becomes byte 00 and is treated like any other data. CR, LF, tabs, and spaces are encoded in their supplied order — there is no newline conversion. Combining sequences remain decomposed; the input "e" followed by U+0301 encodes as 65 CC 81 rather than being rewritten to C3 A9 for U+00E9. The tool does not apply Unicode normalization, case folding, escape parsing, locale-aware rewriting, or any other transformation before encoding.

The one preprocessing step is isolated-surrogate replacement. JavaScript stores strings as UTF-16, so characters outside the BMP occupy two code units called a surrogate pair. If a string contains a high or low surrogate without a matching partner, it is not a Unicode scalar, and the WHATWG TextEncoder standard specifies replacing each such code unit with U+FFFD before encoding. The replacement appears as EF BF BD in hex and is counted visibly in the result panel so you know a round trip will not be exact for that input.

When a Hex Round Trip Will Fail

A round trip is "text → hex → text", and it is exact only when the input is well-formed Unicode that does not begin with U+FEFF. The leading-zero-width-no-break-space caveat exists because some decoders, including the companion Hex to Text Converter, consume a leading EF BB BF as a byte order mark rather than as the character U+FEFF. If your text begins with that character and you decode through a BOM-consuming tool, the recovered text will be missing the first character.

Two other documented failure modes come up often. First, an isolated UTF-16 surrogate cannot survive encoding because it is replaced with U+FFFD; decoding then yields U+FFFD, not your original code unit. Second, decomposed combining sequences stay decomposed through encoding and decode as decomposed sequences — a "well-formed round trip" preserves byte form but not visual normalisation. None of these failure modes are flagged beyond the visible replacement count; you decide whether they matter for your downstream consumer.

Hex is an encoding display, not a one-way function. With the same UTF-8 bytes and a decoder that follows the same policy, you can rebuild the source string. For a reference view of well-known bytes, see the Text to Hex Cheat Sheet — and when you are ready to encode your own string, the Text To HEX tool gives you the exact UTF-8 byte view in continuous, spaced, or 0x-prefixed form.

For a deeper look, see Char Code Lookup for Large Text: Handle Every Code Point.