To turn words into binary, encode each character using UTF-8 and write every resulting byte as exactly eight zero-padded bits separated by a single space, because one letter can occupy anywhere from one to four bytes depending on the character. Converting words to binary is the process of representing human-readable text as a string of 8-bit byte values, where each group of eight 0s and 1s corresponds to a single byte of encoded output. ASCII letters such as A or Z use one byte, accented letters such as é use two, common currency symbols such as € require three, and a supplementary character such as 😀 uses four. The result is therefore not a fixed eight bits per letter — it is eight bits per byte, and the byte count depends entirely on which characters appear in the input string. The Text to Binary Converter encodes the entered string using the browser's standards-based TextEncoder, prints every byte as exactly eight zero-padded bits separated by one ordinary space, and reports the total byte count before you copy it elsewhere. All processing stays in the current tab and is never sent to a remote service.

how to convert words to binary
How to Convert Words to Binary the Right Way

What Converting Words to Binary Actually Means

When people say "convert words to binary," they usually picture each letter turning into its own neat block of eight 0s and 1s. That picture works only when the word is made entirely of plain ASCII characters. The moment the word contains an accent, a punctuation mark outside ASCII, or an emoji, the simple one-letter-equals-one-byte assumption breaks down. Modern text is encoded using UTF-8, a variable-width scheme in which the same character can take one, two, three, or four bytes depending on where its code point sits in the standard. The bytes — not the characters — are what line up neatly on eight-bit boundaries.

UTF-8 is byte-oriented on purpose. Every byte is a self-contained eight-bit unit, and characters that need more than one byte are simply written back-to-back. The decoder walks the bytes from left to right: a byte whose top bit is 0 stands alone as a one-byte character, a byte whose top bits are 110 introduces a two-byte sequence, 1110 introduces three, and 11110 introduces four. That is why a strict byte-exact tool matters — once you lose track of where bytes split, you cannot rebuild the original words. The output of a proper encoder is therefore a stream of bytes, each labeled with exactly eight bits, and the bytes themselves describe the encoded form of the underlying characters.

How UTF-8 Decides How Many Bytes a Word Uses

The number of bytes a word uses is not arbitrary. It is a deterministic function of the code points the word contains, and the rules are spelled out in the WHATWG Encoding Standard and the Unicode Standard. Plain English words made of A through Z stay at one byte per letter. Words that mix in accented letters, currency signs, CJK characters, or emoji will be longer in bytes than the count of letters would suggest, because those characters need two, three, or four bytes each. The tool reports the byte count in the output so you can see the real size of the encoded form.

CharacterCode pointUTF-8 bytesExample output (8-bit groups)
AU+0041101000001
éU+00E9211000011 10101001
€U+20AC311100010 10000010 10101100
😀U+1F600411110000 10011111 10011000 10000000

A mixed-width word such as "café" therefore takes five bytes in UTF-8 — one byte for c, one for a, one for f, and two for é — even though it has four letters. The byte count is what matters when you compare encoded sizes or feed the bytes back into a decoder; the letter count is only useful as a sanity reference. If you want a deeper walkthrough of the byte-by-byte mechanics, see Text to Binary: Why One Character Isn't Always 8 Bits.

How to Convert Words to Binary in Three Steps

The Text to Binary Converter has a small, deliberate interface designed to make byte representation explicit. There is no upload and no third-party processing — the work happens entirely in the current tab. Follow these three steps to encode words as UTF-8 bytes.

  1. Choose the direction. Pick "Text to UTF-8 binary" if you want to encode words, or "UTF-8 binary to text" if you want to decode a binary string back into words.
  2. Enter the input. Type or paste the words into the input field for encoding, or paste the binary — exactly eight 0s and 1s separated by one ordinary space — into the input field for decoding.
  3. Select Convert and review. Trigger the conversion, then check the byte count in the encoding pane or the decoded words in the decoding pane, and copy the result only after you have verified it looks right.

As a worked example, the two-character string "Hi" encodes to the byte sequence 01001000 01101001, which is two bytes and sixteen bits. H corresponds to U+0048, whose eight-bit representation is 01001000; i corresponds to U+0069, whose eight-bit representation is 01101001. Joined with one space, the encoded output reads "01001000 01101001" and the byte count in the output pane shows 2. The encoding itself is described by the browser's TextEncoder, which follows the same UTF-8 rules as the WHATWG standard.

Decoding Binary Back Into Words Safely

Decoding is the inverse operation, and it is just as strict. The decoder requires the input to match an exact regular language: eight binary digits (each either 0 or 1), separated by exactly one ordinary space, with no leading or trailing whitespace. It rejects prefixes such as 0b, commas, tabs, multiple spaces, seven-bit or nine-bit groups, and any character other than 0 or 1. Byte values are never padded, trimmed, guessed, or silently discarded. Decoding accepts up to 180,000 input characters, while encoding accepts up to 20,000 UTF-16 code units from the typed input.

The decoder uses fatal UTF-8 validation through the browser's TextDecoder. A byte sequence that is well grouped can still be invalid text — for example, a lone leading byte without its continuation. Invalid, overlong, surrogate, or out-of-range sequences fail with a clear error rather than being replaced by the Unicode replacement character. This is intentional: it makes the reverse operation useful for checking exact UTF-8 byte strings while avoiding a false round trip. Eight independent golden test cases cover ASCII, a word, two-byte Latin text, a three-byte currency sign, a four-byte emoji, CJK text, a line-feed control byte, and a mixed-width string, with the expected bytes written independently of the implementation. You can read more about the fatal-decoder behavior on MDN's TextDecoder page.

Where Binary Output Goes Wrong in Transport

The output is a plain string of zeros, ones, and spaces, and the spaces are part of the format. Copy or store the result only in systems that preserve ordinary spaces and line content exactly. Chat clients and rich-text editors may collapse runs of spaces or insert line breaks, which would make strict decoding fail even when the bytes themselves are correct. Pasting the output into a word processor or messaging app can therefore change the byte sequence without you noticing, and the decoder will reject the result.

For protocol, source-code, or forensic work, verify the actual byte sequence against the specification and the destination system rather than relying on how a font draws the output. A monospaced font only makes the bytes easier to read on screen — it does not encode any information into the bits. If you need a persistent copy, save the output into a plain-text file or paste it into a code editor that does not autoformat whitespace.

What Binary Text Is Not

This tool is an encoding utility, not encryption. Binary output contains exactly the same information as the original words and offers no secrecy, authentication, integrity, compression, or password protection. Anyone with a decoder can read it. The tool also does not parse numeric binary values, machine instructions, files, images, Base64, hexadecimal, Morse code, or custom legacy character sets. When the real intent is one of those, use a tool that is built for it — for example, a hex converter, a Base64 encoder, a Morse code translator, or a checksum calculator. Picking the right representation is more useful than forcing binary onto a job it was never meant to do.

Used for its actual purpose — inspecting bytes, learning how UTF-8 groups characters, or moving short byte sequences through a system that already expects binary — the Text to Binary Converter gives you a clean, byte-exact, local view of the data without guessing or silently fixing what should be rejected.

If you're weighing options, Bulk UTF-8 Decode in Browser: Hex, Decimal, and Binary covers this in detail.