Decoding UTF-8 in C# means turning a byte array into a string using the .NET runtime's Unicode transformations. The standard path is byte[] bytes then string text = Encoding.UTF8.GetString(bytes), which applies the same Unicode Transformation Format standardized in RFC 3629. C# offers a second path through System.Text.UTF8Encoding, which can be configured to throw on invalid bytes rather than emit U+FFFD. Both paths assume the byte sequence really is UTF-8; if the source was produced as Windows-1252, Shift JIS, GBK, or another legacy encoding, the same byte array will round-trip into a different string. Decoding is therefore a two-step problem: pick a C# method that fits your error policy, then confirm the bytes were UTF-8 in the first place. A local browser decoder lets C# developers paste the byte sequence, compare the decoded text against their code output, and walk away with no data leaving the tab. The verification matters because Encoding.UTF8 will always return a string, even when the underlying bytes cannot represent the source faithfully.

utf8 decode c#
UTF-8 Decode in C#: Bytes to String Without Silent Errors

How UTF-8 Maps Code Points to Bytes

UTF-8 is a variable-width encoding: every valid Unicode scalar value between U+0000 and U+10FFFF is written as one, two, three, or four bytes. ASCII code points from U+0000 through U+007F keep their single-byte form, which is why a file of plain English text looks identical in UTF-8 and ASCII. Code points from U+0080 through U+07FF need two bytes; U+0800 through U+FFFF need three; and supplementary characters from U+10000 through U+10FFFF, including most emoji, need four. The encoding is defined by RFC 3629, and the Unicode Standard 17.0 Chapter 3 lists the same boundaries.

A few well-known anchors help when you eyeball a byte sequence. The dollar sign U+0024 encodes as 0x24. Capital A U+0041 encodes as 0x41. The cent sign U+00A2 crosses the two-byte boundary and encodes as 0xC2 0xA2. The euro sign U+20AC takes three bytes: 0xE2 0x82 0xAC. The grinning face emoji U+1F600 takes four bytes: 0xF0 0x9F 0x98 0x80. The maximum scalar U+10FFFF encodes as 0xF4 0x8F 0xBF 0xBF. The boundary cases 0x7F and 0xE0 0xA0 0x80 mark the transition into multi-byte territory. None of these anchor values are computed by the tool; they are official definitions from the Unicode Consortium and RFC 3629.

Decoding UTF-8 in C# Step by Step

The .NET runtime ships two practical decoders in the System.Text namespace. The first lives on the static Encoding.UTF8 property and is configured for permissive fallback: invalid bytes are replaced by U+FFFD. The second is the UTF8Encoding class, which accepts a constructor flag that turns invalid bytes into thrown exceptions. Both target the same UTF-8 standard, but they answer different questions: Encoding.UTF8 asks "what text should I display?", while the strict UTF8Encoding asks "are these bytes really UTF-8?"

  1. Read the byte array from your source: byte[] bytes = File.ReadAllBytes(path); or byte[] bytes = Convert.FromBase64String(text); depending on where the data came from.
  2. Call the permissive decoder: string text = Encoding.UTF8.GetString(bytes);. This always succeeds and substitutes U+FFFD for any byte that cannot be interpreted, which is the silent behavior you want to avoid in production decoding.
  3. Call the strict decoder when you cannot tolerate replacement characters: var strict = new UTF8Encoding(encoderShouldEmitUTF8Identifier: false, throwOnInvalidBytes: true); string text = strict.GetString(bytes);. A DecoderFallbackException is raised on malformed input instead of producing a quiet U+FFFD.
  4. Inspect the result. If the strict decoder threw, the bytes were not UTF-8. Investigate the upstream source rather than catching the exception and continuing, because swallowing the exception hides the same evidence a replacement character would hide.
  5. Compare against an independent verifier. A local browser decoder accepts the same byte array in hex, decimal, or binary and either produces matching text or refuses to decode. Equal output from two independent decoders is strong evidence the bytes really are UTF-8.

Verifying C# Output in the Browser

A dedicated UTF-8 Encoder / Decoder runs entirely in the current tab, so a C# developer can paste a byte sequence, confirm the result, and close the tab without any byte leaving the browser. The converter uses the browser's TextEncoder and a fatal TextDecoder, so malformed input is rejected instead of quietly replaced, matching the strict behavior of UTF8Encoding(throwOnInvalidBytes: true). The procedure matches the operating steps documented for the converter:

  1. Choose UTF-8 bytes to text on the converter, then pick the notation that matches your C# source data: hexadecimal, decimal, or 8-bit binary.
  2. Enter the byte tokens in the chosen notation. Hex accepts one or two-digit bytes separated by spaces or commas, optional 0x prefixes, or one continuous even-length string. Decimal accepts integer tokens from 0 through 255. Binary requires exactly eight characters per token.
  3. Select the conversion button and read the result panel. The panel reports decoded text and rejects malformed input with an explicit error.
  4. Compare the decoded text with the string returned by Encoding.UTF8.GetString or your strict UTF8Encoding instance. Any difference points to a notation mistake, a stray separator, or non-UTF-8 bytes that need investigation.
  5. Verify a round trip before deleting the source. Encode the decoded text back to the same notation and confirm the byte sequence matches your original input. If it does, the original bytes were already valid UTF-8 and your C# code is safe to trust.

Why Fatal Decoding Beats Replacement Characters

The default Encoding.UTF8 in C# uses a replacement fallback that substitutes U+FFFD for any unparseable byte or sequence. That is convenient when you want display text and do not care about correctness, but it silently destroys evidence. A C# program that receives garbage bytes and prints them out will still produce a string of real Unicode characters; the question of whether those characters faithfully represent the original input is answered no, but the program will not say so.

A fatal decoder flips the question. It rejects overlong encodings such as 0xC0 0xAF, truncated sequences such as 0xE2 0x82 with no continuation byte, isolated continuation bytes such as 0xA2, surrogate encodings such as 0xED 0xA0 0x80, and values above U+10FFFF. The browser converter runs the same fatal mode through TextDecoder('utf-8', {fatal:true}) and raises an explicit error. C# developers get the same property by passing throwOnInvalidBytes: true to UTF8Encoding. The practical difference is that a fatal decoder gives you a reason to investigate the source, while a replacement character gives you a green light to ship corrupted text.

Comparing Hex, Decimal, and Binary Input

The same byte array is shown differently depending on the chosen representation. None of these is a different encoding; they are display layers over the underlying byte sequence, and the tool serializes the same array three ways without touching its meaning.

Code pointMeaningHex bytesDecimal bytesBinary bytes
U+0024Dollar sign243600100100
U+0041Capital A416501000001
U+00A2Cent signC2 A2194 16211000010 10100010
U+20ACEuro signE2 82 AC226 130 17211100010 10000010 10101100
U+1F600Grinning faceF0 9F 98 80240 159 152 12811110000 10011111 10011000 10000000
U+10FFFFMaximum scalarF4 8F BF BF244 143 191 19111110100 10001111 10111111 10111111

Inputs the Decoder Rejects

A few common inputs look almost right but fail strict UTF-8 decoding, and a C# developer should expect the same outcome from a strict UTF8Encoding instance:

  • Overlong encodings such as 0xC0 0xAF for U+002F, where a value is encoded in more bytes than necessary.
  • Truncated multi-byte sequences such as 0xE2 0x82 with no continuation byte following.
  • Isolated continuation bytes such as 0xA2 or 0xBF that appear without a leading byte.
  • Surrogate encodings such as 0xED 0xA0 0x80 representing a UTF-16 surrogate value, which is not a valid Unicode scalar value.
  • Values above U+10FFFF, which exceed the maximum defined scalar.
  • Unpaired UTF-16 surrogate code units inside a JavaScript string passed to TextEncoder on the encoding side, which the tool rejects before encoding to keep a claimed lossless conversion from silently changing input.

The C# rule of thumb is simple: if you cannot prove the bytes were UTF-8, do not let Encoding.UTF8 turn them into a string that looks valid. Use the strict constructor, catch the exception, and identify the original source. The browser converter applies the same rule and gives you a second opinion that runs locally, with input bounded at 200,000 UTF-16 code units for text and 200,000 bytes for decoded notation. Anything larger belongs in a dedicated binary tool, not in an interactive decoder.

If you're weighing options, Base100 Encode Safe: A Byte-Exact Workflow covers this in detail.

If you're weighing options, Base58 Decode on Android Without Installing an App covers this in detail.