Python 3 source files are UTF-8 by default and the str type stores Unicode text in memory, so most everyday Python code already runs on UTF-8. When you need to convert to UTF-8 in Python, the two core operations are str.encode('utf-8') for going from a Unicode string to bytes and bytes.decode('utf-8') for going the other way; this covers string literals, network payloads and any bytes you read in binary mode. The complication starts when the bytes already on disk were produced by something that wasn't UTF-8 — a Windows-1252 export, a UTF-16 little-endian dump, or a file saved with a UTF-8 byte-order mark — because decoding the wrong way either crashes with UnicodeDecodeError or silently replaces characters with the replacement glyph �. Auto-detection libraries such as chardet make a probabilistic guess that can look plausible while still corrupting names, punctuation or currency symbols. For those cases, an explicit source encoding plus a file-based converter such as the UTF-8 Converter produces an auditable, BOM-free UTF-8 file without writing code.

How Python 3 Handles UTF-8 by Default
Since PEP 3120 in 2009, every Python 3 source file is parsed as UTF-8 unless a different encoding is declared in a coding comment. The runtime goes further: str objects hold Unicode code points, the open() built-in defaults to the platform's preferred encoding (which is UTF-8 on modern Linux and macOS), and print() writes through the standard output stream's encoding. That means a string literal like s = "café — naïve façade" already contains the right characters; the question is only how those code points are turned into bytes when you write them somewhere.
Two related points often confuse newcomers. First, "café" in a Python source file is a str, not bytes — it has no encoding of its own until you call encode(). Second, the encoding= argument on open() controls how bytes on disk are turned into str on read (or vice versa on write), not how str is stored internally. Those two facts are why "convert to UTF-8" is really two different jobs in Python: serialising a str you already have, and re-encoding a file that someone else produced in a different encoding.
Converting a Python String to UTF-8 Bytes
For the first job — turning a Unicode string you already have into UTF-8 bytes — the standard library is enough:
- "café".encode("utf-8") produces b'caf\xc3\xa9', the two-byte UTF-8 sequence c3 a9 for é.
- "你好".encode("utf-8") produces three bytes per Chinese character because each one is outside the BMP.
- "🙂".encode("utf-8") produces four bytes because 🙂 is a supplementary character outside the BMP.
If you want a Python source file to declare a different encoding, add a coding comment on the first or second line, for example # -*- coding: latin-1 -*-. This is rarely needed today because UTF-8 covers essentially everything, but it explains why older tutorials mention the directive at all. When writing to disk, prefer the explicit form open(path, "w", encoding="utf-8") over relying on the platform default, especially on Windows where the locale may still default to a legacy code page.
Decoding Bytes Back to a Python String
The reverse operation matters whenever you read a file or a network response in binary mode. open(path, "rb") returns bytes, and you choose how to interpret them:
- data.decode("utf-8") for plain UTF-8.
- data.decode("utf-8-sig") for UTF-8 with a leading BOM, which is consumed silently.
- data.decode("utf-16"), data.decode("utf-16-le") or data.decode("utf-16-be") for UTF-16 files.
- data.decode("cp1252") for Windows-1252.
Pass errors="strict" (the default) so that malformed bytes raise UnicodeDecodeError instead of silently turning into �. The way browsers decode text, including the printable characters in the 0x80–0x9F range (such as the euro sign and curly quotation marks) that ISO-8859-1 leaves undefined, is defined in the WHATWG encoding standard; that is also a common source of "looks right but isn't" surprises when a Windows export claims to be ISO-8859-1.
Why Python's Encoding Auto-Detection Falls Short
The most common way to convert a file to UTF-8 in Python ends up looking like this:
- raw = open(path, "rb").read()
- guess = chardet.detect(raw)["encoding"]
- text = raw.decode(guess)
- open(path, "w", encoding="utf-8").write(text)
That pipeline works well on long, clearly monolingual files where the byte statistics clearly favour one encoding. It is unreliable on short files, on code-switching text, on Windows-1252 vs ISO-8859-1 ambiguity, and on any file where the same byte range is valid under more than one legacy encoding while representing different characters. A euro sign saved as Windows-1252 byte 0x80 can decode to a control character under ISO-8859-1, and a smart-quote byte can decode to nothing readable at all under a different label. The "automatic" step is actually a guess, and a guess that silently passes through decode() with errors="replace" can ship corrupted text downstream without any error to investigate.
Convert a File to UTF-8 Without Writing Python
When you already know the encoding that produced the file — because the originating application, a colleague or reliable metadata told you — the safer pattern is to make that choice explicit and convert the bytes directly. The UTF-8 Converter follows exactly that pattern, entirely in the browser, so the source bytes never leave the machine.
- Identify the source encoding from the producing application or reliable metadata (UTF-8, UTF-16 little-endian, UTF-16 big-endian, or Windows-1252).
- Select that text file — up to 10 MB — and choose the matching encoding from the converter's list.
- Convert and inspect the preview, paying close attention to non-ASCII characters such as names, punctuation, currency symbols and emoji.
- Download the resulting file, which is named with a -utf8 suffix and uses a plain-text UTF-8 media type without a BOM.
- Test the downloaded file in the destination application before replacing any original, and keep the original file until the complete workflow is verified.
UTF-8 mode validates the source with fatal error handling: invalid continuation bytes, truncated sequences and forbidden encodings cause a clear failure rather than replacement characters, and a leading UTF-8 BOM is consumed by the standards-based decoder before re-encoding. UTF-16LE and UTF-16BE selections differ in byte order — the same pair of bytes can become nonsense if endian order is reversed — and a matching BOM is recognised and removed. Windows-1252 mode uses the browser's standards-defined decoder so the printable punctuation in the 0x80–0x9F range is mapped correctly before UTF-8 re-encoding.
Source Encodings the Converter Accepts
| Source encoding | Typical origin | Recognised BOM | Decoder behaviour |
|---|---|---|---|
| UTF-8 | Modern editors, Linux defaults, most web exports | EF BB BF (consumed) | Fatal validation, no silent replacement |
| UTF-16 little-endian | Windows Notepad "Unicode", some Java and .NET tools | FF FE (consumed) | Surrogate pairs combined into one code point |
| UTF-16 big-endian | Some Java tools, network protocols | FE FF (consumed) | Surrogate pairs combined into one code point |
| Windows-1252 | Legacy Windows software, Western European exports mislabelled as ISO-8859-1 | None | 0x80–0x9F printable punctuation mapped before re-encoding |
The output is emitted by the UTF-8 encoder without adding a BOM, and the page reports both source and output byte counts so the expected expansion or contraction is visible. A Windows-1252 euro byte 0x80 becomes the three UTF-8 bytes E2 82 AC; a UTF-16 file typically shrinks because UTF-8 packs ASCII characters into one byte. Different counts are expected and do not by themselves indicate data loss.
Limits and When to Use a Streaming Tool Instead
The converter is bounded to 10 MB because the browser's File and Encoding interfaces operate on an in-memory buffer and the size is checked before reading. For database dumps or multi-gigabyte logs, prefer a trusted streaming utility such as iconv on Linux or Get-Content -Encoding piped to Out-File -Encoding utf8 on Windows, both of which accept explicit source and destination encodings and avoid loading the whole file at once.
The preview helps catch an obviously wrong selection before download, but it may not display every control character or every Unicode normalisation difference. Inspect representative names, punctuation, currency symbols and non-ASCII lines; if the source contains mixed encodings in one file, a single decoder cannot repair it reliably and you will need to split the file first.
Verifying the Output in the Destination App
A clean preview is necessary but not sufficient. Open the downloaded UTF-8 file in the application that will actually consume it — a Python script that reads it with open(path, encoding="utf-8"), a database COPY command, a JSON parser, a CSV importer — and check that names, currency, punctuation and any non-ASCII lines look right end to end. Compare the byte count the page reported against what the destination tool sees; if the importer reports fewer characters than expected, revisit the source encoding choice rather than the output. Only once the full round-trip is verified should the original file be retired.
Related reading: UTF-8 Decode in C#: Bytes to String Without Silent Errors.