Manual text-to-binary conversion means writing each UTF-8 byte of a Unicode string as exactly eight binary digits, separated by a single space. To convert text to binary manually, write down the UTF-8 byte sequence for each character and then expand every byte into eight zero-padded bits. The letter H becomes 01001000 because its ASCII code is 72 and 72 in binary is 1001000. Accented letters such as é take two UTF-8 bytes (11000011 10101001), the euro sign € takes three (11100010 10000010 10101100), and a supplementary character such as the grinning-face emoji 😀 takes four bytes. Writing bits without naming the encoding is the most common reason hand-done conversions disagree: one tool treats é as a single 16-bit Unicode code unit, another treats it as two bytes, and a third silently substitutes the Unicode replacement character when its decoder fails. A byte-exact manual conversion therefore has three non-negotiable rules: pick UTF-8, group the output into strict eight-bit chunks separated by one ordinary space, and never pad, trim, or guess when a sequence is malformed.

Choosing UTF-8 for a Manual Conversion
UTF-8 is the encoding standard that defines how every Unicode code point turns into a specific sequence of one to four bytes. The WHATWG Encoding Standard defines the rules used in every modern browser, and the Unicode Standard defines the code points that UTF-8 maps to those bytes. When you convert text to binary by hand, you are really answering two questions: which encoding did you pick, and which units do you print. The honest answer for a web page, source-code file, or JSON payload is always UTF-8 bytes, because UTF-8 is the on-the-wire format that almost every protocol actually carries.
If you skip the encoding question and write down the numeric value of each character as base-2, you get a result that looks binary but is not bytes. The letter A as a Unicode code point is U+0041, and rendering that scalar value in base-2 gives 1000001, which is only seven digits. Most readers will pad it to 01000001 to keep eight bits per group, but that is a presentation choice, not a network fact. As soon as the input contains é (U+00E9), the code-point-as-base-2 approach gives a 16-bit string that does not match the actual bytes a server would send. The two approaches only agree for ASCII, which is why beginners often think their manual method is correct until they test a non-ASCII word.
Converting a Single Character by Hand
The smallest unit you can practice on is one ASCII letter, because ASCII sits inside UTF-8 as a single byte. Pick H. Look up its ASCII code, which is 72. Convert 72 into base 2 by subtracting the largest powers of two that fit: 72 minus 64 leaves 8, 8 minus 8 leaves 0, and the powers of two you used were 64 and 8, giving bits at positions 6 and 3. Reading those positions from bit 7 down to bit 0 gives 01001000. Zero-pad on the left so the group is exactly eight digits, and the binary for H is 01001000. The whole calculation, written out with repeated division by two:
72 ÷ 2 = 36 remainder 0 36 ÷ 2 = 18 remainder 0 18 ÷ 2 = 9 remainder 0 9 ÷ 2 = 4 remainder 1 4 ÷ 2 = 2 remainder 0 2 ÷ 2 = 1 remainder 0 1 ÷ 2 = 0 remainder 1 Read the remainders bottom-up: 1001000. Zero-pad to 8 bits: 01001000.
That single group is what a strict converter produces for H. Repeat the same divide-by-two procedure for each byte of a longer string and you have a hand-done binary encoding. The mechanical part is easy; the part that breaks is what counts as one byte once the input leaves the ASCII range.
Converting a Full String to Binary
To convert text to binary manually for anything longer than a single letter, follow this exact procedure:
- Decide that UTF-8 is the encoding. Anything else leads to byte counts that do not match what a server, file, or JSON parser actually carries.
- Obtain the UTF-8 byte sequence for the string. You can derive it from an ASCII table for letters and digits, from a UTF-8 chart for accented characters, or from a standards-based encoder.
- For each byte in the sequence, convert the byte's decimal value (0 to 255) into base 2 using repeated division by two.
- Write each result as exactly eight binary digits, adding leading zeros so the group is always eight wide.
- Separate the groups with a single ordinary space. Do not use commas, tabs, the 0b prefix, or multiple spaces.
- Count the groups. For pure ASCII input the byte count equals the character count; for any other input, the byte count is larger and must match the sum of per-character UTF-8 widths.
- If a byte cannot be encoded or the source text contains an invalid surrogate, stop. Manual conversion cannot recover information the source did not contain.
The fifth step is the one most people skip. They join groups with whatever character is handy, then wonder why a decoder refuses the result later. Strict eight-bit groups separated by exactly one space is the format every byte-exact reverse tool accepts, including the Text to Binary Converter.
How Many Bytes Each Character Actually Uses
UTF-8 is a variable-width encoding. The byte count for a character depends on its Unicode range, not on whether it looks like one symbol. The widths below are defined by the Unicode Standard and match what the browser's TextEncoder produces for each class of input.
| Character class | Example | UTF-8 byte count |
|---|---|---|
| ASCII letter or digit | A, 7 | 1 byte |
| Accented Latin | é, ñ | 2 bytes |
| Currency or symbol | € | 3 bytes |
| CJK ideograph | 中, 漢 | 3 bytes |
| Supplementary emoji | 😀, 🎉 | 4 bytes |
The euro sign uses three bytes, which surprises people who expect symbols to share the same width as accented letters. CJK ideographs and the euro sign look very different on screen but happen to fall in the same three-byte range because their code points sit in the U+2000 to U+FFFF band of the Basic Multilingual Plane. Emoji fall outside that plane in the supplementary planes and need four bytes. Treat any manual answer that claims a fixed "one byte per character" as wrong for non-ASCII input.
Where Hand-Done Conversions Disagree
Manual conversions drift from a reference output for five predictable reasons. Each one changes the bytes, not just the formatting.
- Using UTF-16 code units instead of UTF-8 bytes. JavaScript strings are UTF-16, so a quick "charCodeAt" loop produces 16-bit values, not bytes. A supplementary emoji becomes a pair of surrogates, and the resulting binary string no longer matches what the network actually carries.
- Dropping the zero padding. The letter A is 1000001, only seven digits. Without left padding the group is the wrong width, and a strict decoder rejects it. Every group must be exactly eight digits.
- Printing the scalar value in base 2 instead of the bytes. For ASCII this happens to match, which trains the habit. For é the scalar is U+00E9 and the bytes are C3 A9; the binary strings disagree.
- Silently replacing invalid bytes. A decoder configured to use replacement characters will turn malformed input into U+FFFD without complaining. The user thinks the round trip succeeded when in fact the bytes were lost.
- Letting a chat client collapse the spaces. Joining groups with one space is fine in the source, but rich-text editors and chat apps regularly strip or merge runs of whitespace. Once the separator is lost, the strict format is gone and a strict decoder will reject the input.
All five failures share one root cause: a step that was supposed to be exact was treated as a presentation detail. UTF-8 bytes are exact, and the manual procedure has to honor them.
Use the Text to Binary Converter for Strict Output
When the manual procedure is too long, or when you need to verify a hand-done conversion against a reference, use the Text to Binary Converter. The tool encodes Unicode text with the browser's standards-based TextEncoder and prints every byte as exactly eight zero-padded bits separated by one ASCII space, which is the format described above. The reverse mode requires the same strict format and uses a fatal UTF-8 decoder, so a malformed byte sequence produces a clear error instead of being silently replaced by U+FFFD. Eight independent golden cases, covering ASCII, a word, two-byte Latin text, a three-byte currency sign, a four-byte emoji, CJK text, a line-feed control byte, and a mixed-width string, are checked against expected bytes that were written independently of the implementation.
A few practical limits apply. Encoding accepts up to 20,000 UTF-16 code units, and decoding accepts up to 180,000 input characters. The decoder rejects prefixes such as 0b, commas, tabs, multiple spaces, seven-bit groups, nine-bit groups, leading or trailing whitespace, and any character other than 0 or 1. Values are never padded, trimmed, guessed, or silently discarded. Inputs and outputs stay in the current tab and are not uploaded, so the tool is also appropriate for the same privacy-sensitive strings that justify hand conversion in the first place. MDN documents the fatal TextDecoder option that makes the decoder fail loudly on invalid UTF-8, which is the behaviour behind the strict format rule.
Binary output contains the same information as the original text. It is a reversible representation, not encryption, so anyone who receives the bytes can read the original characters. Once the conversion is done, copy or store the result only in systems that preserve ordinary spaces and line content exactly, such as a plain-text file or a code block in a Markdown document. Chat clients and rich-text editors may collapse spaces or insert line breaks, which would make strict decoding fail later. For protocol, source-code, or forensic work, verify the actual byte sequence against the destination system rather than relying on how a font draws the output.
Related reading: Text to Hex Example: What UTF-8 Bytes Look Like in Practice.
Related reading: Decode UTF-8 Bytes Without Silent Replacement Characters.