To convert text to binary numbers means turning every character of a Unicode string into a sequence of bytes, then writing each byte as exactly eight binary digits separated by spaces. In UTF-8, the standard text encoding used on the web, ASCII letters become one byte, accented Latin letters become two, the euro sign and most CJK characters become three, and supplementary-plane characters such as emoji become four. The Text to Binary Converter performs that mapping locally in your browser using the WHATWG TextEncoder, prints each byte as eight zero-padded bits, and never collapses, trims, or guesses values. If you paste "Hi", the tool outputs 01001000 01101001 — two bytes, sixteen bits total — because the letter H is ASCII 72 (binary 01001000) and the letter i is ASCII 105 (binary 01101001). The same logic applies to longer strings: byte count grows with the encoded characters, not with the number of visible letters, which is why understanding UTF-8 widths is essential before you trust any binary output.

What "Text to Binary Numbers" Really Means
The phrase sounds simple, but it hides two separate steps. First, every character is mapped to a numeric code by a character encoding; in modern systems that encoding is almost always UTF-8, the same byte format used in HTML, JSON, source files, and network protocols. Second, each byte — a number between 0 and 255 — is written in base 2 as eight digits of 0 and 1, padded on the left with zeros so every group is the same length. The output you see is therefore not a single integer per character; it is a stream of byte values, and the number of bytes per character varies.
A common misconception is that one character equals eight bits. That rule only holds for the ASCII range (code points 0 through 127). Outside ASCII, UTF-8 uses continuation bytes to express larger code points, and a single character can take two, three, or four bytes. The Text to Binary Converter prints encoded bytes, not JavaScript UTF-16 code units, glyph pixels, abstract bits without grouping, or a base-2 rendering of each Unicode scalar value. That distinction matters whenever you compare output from two tools: if a tool claims to convert "characters" to binary and produces a single eight-bit group per glyph for emoji, it is doing something other than byte-exact UTF-8.
UTF-8 Byte Widths at a Glance
The table below lists the four UTF-8 byte widths and a representative character for each. These widths are defined by the UTF-8 standard itself, not by any particular tool.
| Character type | UTF-8 bytes | Example | Byte sequence (hex) |
|---|---|---|---|
| ASCII letter or digit | 1 | A (U+0041) | 41 |
| Accented Latin letter | 2 | é (U+00E9) | C3 A9 |
| Currency sign, CJK, BMP symbol | 3 | € (U+20AC) | E2 82 AC |
| Supplementary-plane emoji or symbol | 4 | 😀 (U+1F600) | F0 9F 98 80 |
A worked check for the simplest case: the letter A has the ASCII decimal code 65. Converting 65 to base 2 with zero-padding to eight digits gives 01000001 — the same byte you see when an ASCII byte with value 0x41 is written in binary. This single-byte rule is what makes early tutorials work cleanly, but it stops being true the moment you leave the ASCII range.
Convert Text to Binary Numbers in Three Steps
The fastest path from a string to a clean byte stream is to use the online tool rather than writing a script. The following steps match the converter's verified operating procedure.
- Choose a mode. Pick Text to UTF-8 binary when you are encoding characters, or UTF-8 binary to text when you are decoding an existing binary string back to characters.
- Enter your input. In encode mode, type or paste any Unicode text. In decode mode, paste exact 8-bit byte groups separated by a single ordinary space; no commas, tabs, prefixes, or extra whitespace.
- Select Convert and verify. Read the byte count or decoded text shown next to the output, confirm it matches your expectation, and only then copy the result into the system that needs it.
All processing happens in the current browser tab. Your input and output are not uploaded to any server, which is convenient when the text is sensitive but still valuable when you simply want a fast round trip during development.
Reading the Output and Counting Bytes
Encoded output is a single line of space-separated byte groups, with one group per byte. Counting the groups gives you the encoded byte count for your string, which is the number you need when you are sizing buffers, comparing file sizes, or matching a protocol specification. A string of n ASCII characters produces exactly n groups; a string that contains one emoji produces n plus four groups, because the emoji adds four bytes — one lead byte plus three continuation bytes.
The converter pads every group to exactly eight digits. That padding is not decorative — it is what lets the decoder parse each group unambiguously as a byte. If you ever see a tool that emits 7-bit groups, 9-bit groups, or unpadded groups of varying length, treat that output as a different format and do not feed it into a strict 8-bit decoder. For a deeper look at the reverse direction, the guide on how binary to text conversion works in plain English covers the same byte-stream logic from the other side.
Why Strict Eight-Bit Grouping Matters
Binary text is ambiguous unless the encoding and grouping rule are both stated. A group of eight bits can mean a byte (0–255), but a group of seven bits is the ASCII tradition, and JavaScript strings are internally UTF-16 code units (0–65535), not bytes. The Text to Binary Converter chooses bytes and refuses to silently reinterpret your input: it rejects prefixes such as 0b, commas, tabs, multiple spaces, seven- or nine-bit groups, leading or trailing whitespace, and any character outside 0 and 1. Values are never padded, trimmed, guessed, or dropped.
The decoder side is just as strict. It uses fatal UTF-8 validation, which means a byte sequence that is structurally well grouped can still be invalid text — for example, a lone leading byte without its continuation. Instead of substituting the Unicode replacement character (U+FFFD) and continuing, the tool reports a clear error. That behaviour protects round-trip integrity for protocol, source-code, and forensic work, where a silent replacement would corrupt the meaning of the byte stream.
Limits, Errors, and What the Tool Does Not Do
The converter accepts up to 20,000 UTF-16 code units on the encoding side and up to 180,000 input characters on the decoding side. Inputs that exceed those budgets are rejected before any conversion runs. Invalid grouping, repeated spaces, incomplete UTF-8 sequences, and out-of-range or surrogate values also fail with an explicit message rather than a replacement glyph.
It is worth being clear about what the tool is not. It is not encryption: the binary output contains the same information as the original text and offers no secrecy, authentication, integrity, compression, or password protection. It does not parse numeric binary values, machine instructions, files, images, Base64, hexadecimal, Morse code, or custom legacy character sets; if any of those is the real intent, use the neighbouring tool that targets that format. Finally, copy or store the result only in systems that preserve ordinary spaces and line content exactly; chat clients and rich-text editors may collapse spaces or insert line breaks, which would make strict decoding fail even when the bytes themselves are correct.
If you're weighing options, Text to Hex Explained: How UTF-8 Bytes Become Hex Pairs covers this in detail.