Base64 解碼失敗幾乎都來自下列五個根本原因之一:錯誤的字母表(URL-safe 與 standard)、遺漏或被移除的填充字元、字串中夾帶的空白字元、輸出端的字元集混淆,或輸入根本就不是 Base64。「tips and common mistakes」這個搜尋詞來自那些至少遇過其中一個問題、想在浪費下一個小時胡亂嘗試之前先有一份乾淨檢查清單的人。一個可靠的 Base64 解碼器應該嚴格遵循 RFC 4648 字母表(A–Z、a–z、0–9,以及 '+' 與 '/'),強制執行 '=' 填充規則,拒絕該字母表以外的字元而非默默回傳垃圾,並透過嚴格的 UTF-8 驗證器解碼所得的位元組,讓你看到真正的錯誤而不是亂碼文字。大多數瀏覽器的程式碼片段都會在 Unicode 這一關栽跟斗:內建的 btoa() 一看到 café、你好 或 😀 就會立刻拋出例外,因為它只接受 0–255 的字碼點。正確的作法是在 Base64 編碼前先轉成 UTF-8 位元組,並在解碼後讓位元組流再跑一次嚴格的 UTF-8 解碼器。

The Five Recurring Base64 Decode Mistakes
After watching thousands of decoding attempts, the same handful of mistakes come up over and over. Treating them as a checklist is faster than debugging case by case.
- Wrong alphabet variant. Standard Base64 uses '+' and '/'; URL-safe Base64 (also called base64url) replaces them with '-' and '_'. If your string came from a JWT, an OAuth token, or a URL parameter, it is probably URL-safe. A strict standard decoder will throw "Invalid character" the moment it hits '-'.
- Padding removed. Trailing '=' signs are not decoration. Strip them and a strict decoder will complain about an output length that is not a multiple of four.
- Whitespace in the input. Newlines, spaces, and tabs are routinely inserted when people paste from logs, emails, or terminal output. Some decoders tolerate them; many do not.
- Output decoded with the wrong character set. The bytes are correct, but you are reading them as Latin-1 or Windows-1252 instead of UTF-8. This is the most common cause of "it decoded but the result is gibberish."
- Input that was never Base64. Random JSON, a partial HTML snippet, or a hash printed without context can all look alphabet-ish. A strict decoder flags it; a lenient one returns garbage.
Padding Is Structural, Not Optional
Base64 groups three input bytes (24 bits) into four 6-bit output characters. When the input length is not a multiple of three, the final group has fewer than 24 bits and the output is padded with one or two '=' signs so the encoded length stays a multiple of four — exactly as RFC 4648 §4 requires. That is why 'f' (1 byte) becomes 'Zg==' and 'foobar' (6 bytes, already a multiple of three) becomes 'Zm9vYmFy' with no padding at all.
Most decoders accept missing padding as a convenience, but if your pipeline compares output lengths, builds HMACs, or feeds the result into another RFC 4648–compliant system, leaving the '=' off can break the next step. A useful rule: if your Base64 string's length is not a multiple of four, either it is missing padding or it contains a stray character that needs to be removed.
The Hidden Whitespace Trap
Logs, email clients, and terminal output routinely wrap long Base64 lines at 60 or 76 characters and insert a newline every few kilobytes (this is the MIME line-wrapping convention). If you copy from one of these sources and paste the whole block, the newlines travel with the string. Some decoders ignore them, but atob in browsers and many server-side libraries do not.
The cleanest fix is to strip every byte that is not in the Base64 alphabet before decoding — that includes spaces, tabs, '\r', and '\n'. Visually inspecting the string usually reveals a stray newline where the wrap broke; if the string is too long to scan, search it for the literal sequence '\n' or a carriage return. The same applies to spaces accidentally added by formatted chat messages, which is why a fast first step in any decode attempt is "remove everything outside A–Z, a–z, 0–9, +, /, and =."
When the Output Looks Like Garbage
A successful decode that produces nonsense text almost always points to a character-set mismatch on the output side, not an encoding bug. Base64 is a transport format: it does not know or care whether the underlying bytes are ASCII, UTF-8, ISO-8859-1, or UTF-16. The decoder returns bytes; how those bytes are interpreted is a separate step.
The most common modern case is UTF-8. If the original text contained "café", "你好", or "😀", the encoder had to first convert those characters to multi-byte UTF-8 sequences — 'é' becomes two bytes (0xC3 0xA9), "你" becomes three bytes (0xE4 0xBD 0xA0), and 😀 becomes four bytes (0xF0 0x9F 0x98 0x80). If the decoder hands those bytes back to a Latin-1 reader, you get a string of accented characters or replacement symbols instead of the original.
The browser's built-in btoa() cannot even start this conversation — it throws InvalidCharacterError on any character above code point 255. The reason is that btoa() treats each JavaScript character as one byte, so anything outside the Latin-1 range is out of bounds before encoding begins. A correct pipeline first calls TextEncoder to turn the string into UTF-8 bytes, then applies Base64 to those bytes. On the way back, a strict UTF-8 decoder with fatal: true rejects malformed sequences instead of inserting the replacement character (U+FFFD) and silently corrupting your data.
How to Decode Base64 the Right Way
- Open the Base64 Encode / Decode tool in your browser. Choose the Decode direction so the input box accepts Base64 and the output box shows the recovered text.
- Paste the Base64 string into the input box. As you type or paste, the result updates instantly in the output box below — no submit button to press, nothing sent to a server.
- Scan the output for two failure modes: an error message, or text that looks like é or ’ instead of the expected accents and emoji. An error means the input itself is malformed (bad padding, stray character, wrong alphabet variant). Garbled text means the bytes are right but you are reading them with the wrong character set on the receiving side.
- If you need to go the other direction, click Swap to feed the decoded text straight back through the encoder. This is a fast sanity check: encode → decode should round-trip your original input exactly, including emoji and accents.
- Click Copy to grab the result. For longer pieces, paste the round-trip output into a diff against the original to confirm nothing was lost.
Diagnosing Decoding Errors by the Message You Get
Different symptoms usually point to different mistakes, so a quick mental table saves time.
| Symptom | Most likely cause | First thing to try |
|---|---|---|
| "Invalid character" thrown immediately | URL-safe alphabet (- or _) fed to standard decoder, or whitespace mixed in | Replace - → + and _ → /, then strip spaces, tabs, newlines |
| "Invalid character" only at the end | Padding '=' stripped | Append '=' until length is a multiple of four |
| Decodes without error but output is garbled | Wrong output character set (Latin-1 vs UTF-8) | Decode as UTF-8 bytes, not Latin-1 |
| Output is exactly 0 or 1 characters long | Input was not Base64 at all (random JSON, hash, etc.) | Confirm the source — log, token, MIME attachment? |
| Tool throws on emoji or Chinese in the original | Decoding straight from string without TextEncoder | Encode the original text to UTF-8 bytes first |
A Worked Example: How "Hi" Becomes "SGk="
One concrete walkthrough is enough to make the rules stick. Take the two characters "Hi" — that is 2 bytes (16 bits), which is not a multiple of three, so the encoded output will be three Base64 characters plus one '=' pad.
- Bytes: 'H' = 0x48 = 01001000, 'i' = 0x69 = 01101001.
- Concatenate the bits: 0100100001101001.
- Split into 6-bit groups from the left: 010010 | 000110 | 100100 (the last group has only 4 bits, padded with two zeros to reach 6).
- Map each 6-bit value to the Base64 alphabet: 18 → 'S', 6 → 'G', 36 → 'k'.
- Add one '=' pad because the input is 2 bytes (one short of a multiple of three): SGk=.
If you drop the '=', a strict decoder will reject "SGk" because its length is not a multiple of four. If you try to decode "SGk=" through the browser's atob() after first stuffing UTF-8 bytes back through it, the round-trip still works for "Hi" because both characters are ASCII — but try the same trick with "你好" and atob() throws before it even starts.
Common Places Base64 Strings Hide
Knowing where Base64 strings come from makes the mistakes easier to spot in context. The most common sources are:
- JSON Web Tokens. A JWT has three dot-separated segments; the first two are Base64URL (note: URL-safe, not standard) and the third is the signature. Decoding the middle segment gives you the claims payload; the signature segment is meant to be verified, not read.
- HTTP Basic auth headers. The Authorization: Basic … header is literally "username:password" encoded to Base64 — and that is the only "encryption" it has. Anyone with the header can decode it.
- data: URIs. Inline images and fonts in HTML/CSS are Base64-encoded blobs. They decode cleanly with the right tool and almost never contain whitespace unless someone has broken the URI by hand.
- Email attachments (MIME). Email bodies wrap Base64 at 76 characters with CRLF line endings. Strip the line breaks first.
- API payloads. Some APIs serialize binary blobs as JSON strings of Base64. Decode those bytes as whatever format the API documents (usually UTF-8 text or a raw byte stream).
For a longer reference on the alphabet, padding lengths, and standard versus URL-safe variants, see the Base64 decode cheat sheet — it pairs well with the error-symptom table above.
One last note on privacy: Base64 is a reversible encoding, not a cipher. It exists to move arbitrary bytes through channels that only accept text, and anyone who has the string can decode it without a key. Treat it as transport, not as protection. If you need confidentiality, pair it with real encryption such as AES-256-GCM or use a signed token format; the encoded bytes themselves will still be readable.
For a deeper look, see Convert Base64 to Hex for Large Text Without Losing Bytes.