How to Prepare Your Python Code for the Python 3.15 UTF-8 Default and Fix Windows Encoding Bugs
Python 3.15 makes UTF-8 the default for open(), pipes and stdio. This Windows tutorial uses Python 3.13 and 3.15 RC3 to show which bugs the change fixes, which it exposes, and how to make your code...
Python 3.15 changes what open() does when you do not say which text encoding to use. On Windows, that call has fallen back to the machine’s legacy code page, which on my test machine is cp1252. In 3.15 it falls back to UTF-8 instead. The release notes say it directly: “Python now uses UTF-8 as the default encoding, independent of the system’s environment.” The change cures a famous class of Windows bugs, and it also breaks any program that quietly depended on the old default.
Table Of Content
- The ideas you need first
- Prerequisites
- Step 1: See one character as four byte sequences
- Step 2: Ask Python what it will do by default
- Step 3: Create test files that come from different sources
- Step 4: Write the program with implicit defaults and run it under both modes
- Reading the old default (the first and last blocks)
- Reading the new default (the middle two blocks)
- One more place the old default bites: printing
- Step 5: Preview the new default before you upgrade
- Step 6: Find every implicit default
- Tool 1: let Python warn you at runtime
- Tool 2: scan the source code without running it
- Tool 3: let Ruff check it in CI
- Putting the three tools together
- Step 7: Fix each call by naming the encoding of its source
- Step 8: Make the test suite prove it
- Common mistakes and gotchas
- Check the whole thing end to end
- Where to go next
In this tutorial you will build a small report tool that reads and writes text from four different sources, run it under the old default and the new default on a real Windows 11 machine, and watch one setting turn the same code into correct output, garbled text, or a crash. Then you will find every implicit default with three different tools, fix each call by naming the encoding of its source, and add a test suite that proves the fixed code behaves identically under both defaults. You can start today without installing Python 3.15, because one environment variable previews the new behavior on any recent Python.
The ideas you need first
A file on disk holds bytes, not letters. An encoding is the rulebook that maps letters to bytes and back. Reading bytes with the wrong rulebook produces mojibake, the garbled text you get when, for example, é turns into é. When your code calls open("notes.txt") without an encoding= argument, Python has to guess the rulebook. That guess is the default encoding.
Before 3.15, the guess on Windows came from the locale encoding, the code page configured for the system. According to Microsoft’s documentation, Windows code pages are commonly called ANSI code pages, and code page 1252 is commonly used for English and other Western European languages. Console programs use a different one, the OEM code page. The same page says of OEM code pages: “These code pages were originally used for MS-DOS and are still used for console applications.” It adds: “The usual OEM code page for English is code page 437.” So one Windows machine can hold text in four encodings at once: UTF-8 from Linux, git and the web; ANSI from older tools; OEM from console programs; and UTF-16 from some Microsoft tools.
Python 3.15 turns on UTF-8 mode by default. PEP 686 states the goal: “With this change, Python consistently uses UTF-8 for default encoding of files, stdio, and pipes.” Python source files were already read as UTF-8 (“The default encoding of Python source files is UTF-8”, says the same PEP), so the change brings your data files in line with your source files.
Prerequisites
- Windows 10 or 11 and a PowerShell terminal. I used Windows 11 (version 10.0.26100), where the ANSI code page is cp1252 and the OEM code page is 437, the usual setup for English. Other regional settings use other code pages, so your exact bytes may differ, but the lessons are the same. Step 2 shows how to read your own settings.
- Python 3.11 or newer. The function
locale.getencoding(), which this tutorial uses, was added in 3.11. I used 3.13.14. - Optional: a Python 3.15 build. I used 3.15.0rc3, the newest build in python.org’s 3.15.0 download folder on October 5, 2026. PEP 790 lists the final release as expected on Friday, 2026-10-09. You do not need an installer: the zip file runs in place.
- pytest and Ruff, installed with pip. I used pytest 9.1.1 and Ruff 0.16.10.
- Basic Python: functions, files, and reading a traceback. Everything else is explained as it appears.
Set up a working folder with a virtual environment and the optional 3.15 build. Putting the virtual environment on the search path for the current session avoids the PowerShell script-execution policy that blocks Activate.ps1 on many machines:
mkdir encoding-lab
cd encoding-lab
python -m venv .venv
$env:PATH = "$PWD\.venv\Scripts;$env:PATH"
python -m pip install pytest ruff
curl.exe -L -o py315.zip https://www.python.org/ftp/python/3.15.0/python-3.15.0rc3-amd64.zip
Expand-Archive py315.zip -DestinationPath py315
.\py315\python.exe -m ensurepip --upgrade
.\py315\python.exe -m pip install pytest
.\py315\python.exe --version
The last command should print Python 3.15.0rc3. Recent versions of Windows include curl.exe; type the .exe, because in Windows PowerShell 5.1 plain curl is an alias for a different command. Every file below goes into encoding-lab, and the first line of each code block (a comment such as # matrix.py) is the file name.
Step 1: See one character as four byte sequences
The quickest way to understand every bug in this tutorial is to look at one character, é, in four encodings. Create same_char.py:
# same_char.py
import sys
sys.stdout.reconfigure(encoding="utf-8") # this demo prints odd characters, so pin its own output
text = "é"
print(f"{'encoding':<10} {'bytes':<8} {'read as cp1252':<16} read as UTF-8")
for name in ("utf-8", "cp1252", "cp437", "utf-16-le"):
data = text.encode(name)
as_cp1252 = data.decode("cp1252", errors="replace")
as_utf8 = data.decode("utf-8", errors="replace")
print(f"{name:<10} {data.hex(' '):<8} {as_cp1252!r:<16} {as_utf8!r}")
The script encodes é four ways, then decodes each result with the wrong rulebooks on purpose (errors="replace" makes Python substitute the replacement character � instead of raising an error). The line after the imports pins the script’s own output to UTF-8, so it can print any character on any Python. Run it:
python same_char.py
encoding bytes read as cp1252 read as UTF-8
utf-8 c3 a9 'é' 'é'
cp1252 e9 'é' '�'
cp437 82 '‚' '�'
utf-16-le e9 00 'é\x00' '�\x00'
Read the rows one at a time. UTF-8 stores é as the two bytes c3 a9, and a program that reads those bytes as cp1252 sees two characters, Ã and ©. That is the classic mojibake. The cp1252 row is the reverse: one byte, e9, which UTF-8 cannot accept on its own, so a strict UTF-8 reader refuses it. The cp437 row is what a console program writes for é (byte 82); read as cp1252 it becomes the low quotation mark ‚. The last row is UTF-16, where each common character takes two bytes and an ASCII letter carries a zero byte beside it.
Keep this table in mind. Every bug below is one of these rows read through the wrong column.
Step 2: Ask Python what it will do by default
Create settings.py. It prints the encoding settings that matter, including what open() uses when you give it no encoding= argument:
# settings.py
import locale
import os
import sys
with open(os.devnull) as devnull:
open_default = devnull.encoding
with open(os.devnull, encoding="locale") as devnull:
open_locale = devnull.encoding
rows = [
("Python", sys.version.split()[0]),
("locale.getencoding()", locale.getencoding()),
("locale.getpreferredencoding(False)", locale.getpreferredencoding(False)),
("sys.flags.utf8_mode", sys.flags.utf8_mode),
("open(...).encoding, no encoding=", open_default),
('open(..., encoding="locale").encoding', open_locale),
("sys.getfilesystemencoding()", sys.getfilesystemencoding()),
("sys.getdefaultencoding()", sys.getdefaultencoding()),
]
width = max(len(name) for name, _ in rows)
for name, value in rows:
print(f"{name:<{width}} {value}")
You also need a way to run any script under both defaults, and under both Python versions, without your shell environment interfering. Create matrix.py:
# matrix.py
import os
import subprocess
import sys
from pathlib import Path
sys.stdout.reconfigure(encoding="utf-8")
args = sys.argv[1:]
pythons = []
while args and args[0] == "--python":
pythons.append(args[1])
args = args[2:]
if not pythons:
pythons = [sys.executable]
candidate = Path(__file__).parent / "py315" / "python.exe"
if candidate.exists():
pythons.append(str(candidate))
def clean_env(utf8=None):
env = {k: v for k, v in os.environ.items() if not k.upper().startswith("PYTHON")}
if utf8 is not None:
env["PYTHONUTF8"] = utf8
return env
def ask(exe, code):
out = subprocess.run([exe, "-c", code], env=clean_env(), capture_output=True)
return out.stdout.decode("ascii", "replace").strip()
for exe in pythons:
version = ask(exe, "import sys; print(sys.version.split()[0])")
default_on = ask(exe, "import sys; print(sys.flags.utf8_mode)") == "1"
flip = "0" if default_on else "1"
for label, utf8, mode_on in [("default", None, default_on), (f"PYTHONUTF8={flip}", flip, not default_on)]:
proc = subprocess.run([exe, *args], env=clean_env(utf8), capture_output=True)
print(f"=== Python {version}, {label} (UTF-8 mode {'on' if mode_on else 'off'})")
out = proc.stdout.decode("utf-8", "backslashreplace").replace("\r\n", "\n").rstrip()
if out:
print(out)
for line in proc.stderr.decode("utf-8", "backslashreplace").strip().splitlines():
print("[stderr]", line)
if proc.returncode:
print(f"[exit {proc.returncode}]")
print()
This helper does four things. It finds the interpreter that runs it, plus a Python 3.15 interpreter in py315\python.exe if that folder exists. For each interpreter it asks, with a clean environment, whether UTF-8 mode is on by default. It then runs your command twice: once with the interpreter’s own default, and once with PYTHONUTF8 set to the opposite value. And it first removes every environment variable that starts with PYTHON from the child process, because those variables change the result. While preparing this tutorial I found PYTHONUTF8=1 and PYTHONIOENCODING=utf-8 already set in my own shell. With PYTHONUTF8=1 set, Python 3.13 already behaves like 3.15, so the old-default bugs below never show up. Check yours with Get-ChildItem Env:PYTHON*.
python matrix.py settings.py
The output has one block per run. Here it is reorganized as a table, one column per run:
| Setting | 3.13.14 default | 3.13.14 with PYTHONUTF8=1 | 3.15.0rc3 default | 3.15.0rc3 with PYTHONUTF8=0 |
|---|---|---|---|---|
sys.flags.utf8_mode |
0 | 1 | 1 | 0 |
locale.getencoding() |
cp1252 | cp1252 | cp1252 | cp1252 |
locale.getpreferredencoding(False) |
cp1252 | utf-8 | utf-8 | cp1252 |
open(...).encoding, no encoding= |
cp1252 | utf-8 | utf-8 | cp1252 |
open(..., encoding="locale").encoding |
cp1252 | cp1252 | cp1252 | cp1252 |
sys.getfilesystemencoding() |
utf-8 | utf-8 | utf-8 | utf-8 |
sys.getdefaultencoding() |
utf-8 | utf-8 | utf-8 | utf-8 |
Read it by columns. The UTF-8 mode row decides what open() does, and the Python version does not matter: 3.13 with PYTHONUTF8=1 matches the 3.15 default, and 3.15 with PYTHONUTF8=0 matches the 3.13 default. The Python docs confirm both halves: “Changed in version 3.15: Python UTF-8 mode is now enabled by default (PEP 686). It may be disabled by setting PYTHONUTF8=0 as an environment variable or by using the -X utf8=0 command line option.”
Three rows deserve a closer look. First, locale.getencoding() says cp1252 in all four runs because, in the docs’ words, “This function is similar to getpreferredencoding(False) except this function ignores the Python UTF-8 Mode.” It tells you what the machine is configured for, whatever Python decides to do. Second, open(..., encoding="locale") still gives cp1252 under UTF-8 mode, which makes it the explicit way to ask for the old behavior. Third, sys.getdefaultencoding() says utf-8 everywhere, even on the old default. It describes how Python converts between str and bytes in memory, not how it reads files, so it is not a useful check here.
If locale.getencoding() on your machine already says utf-8, the two defaults agree and the old-default bugs below will not reproduce for you. Microsoft documents that the ANSI code page can be configured for UTF-8 in its guide to the UTF-8 code page.
Step 3: Create test files that come from different sources
Real programs read files written by other programs. The next script creates three small files that imitate three sources you will meet on Windows. Notice that this script passes an explicit encoding every time it writes text; that is the habit this tutorial builds. Create make_fixtures.py:
# make_fixtures.py
import json
import subprocess
from pathlib import Path
customers = [
{"name": "Zoë Müller", "city": "Zürich"},
{"name": "Renée O’Brien", "city": "Cork"},
{"name": "Åsa Lindqvist", "city": "Malmö"},
]
# 1. What a Linux CI job or a web API writes: UTF-8.
Path("customers.json").write_text(json.dumps(customers, ensure_ascii=False, indent=2), encoding="utf-8")
# 2. What a legacy Windows tool writes when told to save as "ANSI": cp1252 on this machine.
Path("shop.ini").write_text("[shop]\nname = Café Délice\nowner = José Núñez\n", encoding="cp1252")
# 3. What Windows PowerShell 5.1 writes for the > operator: UTF-16 LE with a byte order mark.
subprocess.run(["powershell", "-NoProfile", "-Command", '"Zoë Müller" > ps_export.txt'], check=True)
for name in ("customers.json", "shop.ini", "ps_export.txt"):
data = Path(name).read_bytes()
offset = next(i for i, b in enumerate(data) if b >= 0x80)
print(f"{name:<15} {len(data):>3} bytes, first non-ASCII byte at {offset:>2}: {data[offset:offset + 6].hex(' ')}")
python make_fixtures.py
customers.json 194 bytes, first non-ASCII byte at 23: c3 ab 20 4d c3 bc
shop.ini 48 bytes, first non-ASCII byte at 18: e9 20 44 e9 6c 69
ps_export.txt 26 bytes, first non-ASCII byte at 0: ff fe 5a 00 6f 00
Each line prints the first byte above 127, the first byte that is not plain ASCII, and the bytes after it. In customers.json it is c3 ab, the UTF-8 form of ë. In shop.ini it is a lone e9, the cp1252 form of é. In ps_export.txt it is ff fe at offset 0, a byte order mark (BOM) that announces UTF-16 little-endian, followed by 5a 00 6f 00, the letters Z and o, each padded with a zero byte.
The third file comes from powershell, which on my machine is Windows PowerShell 5.1. Microsoft’s documentation says that redirecting with the > operator is “functionally equivalent to piping to Out-File with no extra parameters”, and the bytes above show what that cmdlet writes by default: UTF-16.
Step 4: Write the program with implicit defaults and run it under both modes
Here is a small report tool written the way most Python code is written, with no encoding= anywhere. Create reportkit_v1.py:
# reportkit_v1.py
import configparser
import json
import subprocess
from pathlib import Path
def load_customers(path):
with open(path) as f:
return json.load(f)
def load_shop(path):
parser = configparser.ConfigParser()
parser.read(path)
return dict(parser["shop"])
def write_report(path, customers):
with open(path, "w") as f:
for c in customers:
f.write(f"✓ {c['name']} ({c['city']})\n")
def read_report(path):
return Path(path).read_text()
def native_echo(text):
result = subprocess.run(["cmd", "/c", "echo", text], capture_output=True, text=True)
return result.stdout.strip()
def load_ps_export(path):
with open(path) as f:
return f.read()
It has five jobs. load_customers reads JSON that a Linux job produced. load_shop reads a legacy settings file with configparser. write_report and read_report write a report that uses a check mark and read it back. native_echo runs a console command and captures its output; the subprocess docs say the pipes are “opened in text mode using the specified encoding and errors or the io.TextIOWrapper default”, which means the default encoding again. load_ps_export reads the PowerShell file.
Now a driver that calls each function inside try/except, so one failure does not hide the others. Create demo.py:
# demo.py
import importlib
import sys
from pathlib import Path
sys.stdout.reconfigure(encoding="utf-8") # let the demo print any character, whatever the default is
kit = importlib.import_module(sys.argv[1])
report = Path("report.txt")
def attempt(label, func):
try:
result = func()
except Exception as exc:
print(f"{label:<10} FAIL {type(exc).__name__}: {exc}")
else:
print(f"{label:<10} OK {result!r}")
def roundtrip():
report.unlink(missing_ok=True)
kit.write_report(report, [{"name": "Zoë Müller", "city": "Zürich"}])
return kit.read_report(report)
attempt("customers", lambda: [c["name"] for c in kit.load_customers("customers.json")])
attempt("shop", lambda: kit.load_shop("shop.ini"))
attempt("report", roundtrip)
print(f"report.txt is {report.stat().st_size} bytes after that attempt")
attempt("native", lambda: kit.native_echo("café"))
attempt("ps_export", lambda: kit.load_ps_export("ps_export.txt").strip())
The module name is an argument, so you can point the same driver at the fixed version later. Run it under all four combinations:
python matrix.py demo.py reportkit_v1
=== Python 3.13.14, default (UTF-8 mode off)
customers OK ['Zoë Müller', 'Renée O’Brien', 'Åsa Lindqvist']
shop OK {'name': 'Café Délice', 'owner': 'José Núñez'}
report FAIL UnicodeEncodeError: 'charmap' codec can't encode character '\u2713' in position 0: character maps to <undefined>
report.txt is 0 bytes after that attempt
native OK 'caf‚'
ps_export OK 'ÿþZ\x00o\x00ë\x00 \x00M\x00ü\x00l\x00l\x00e\x00r\x00\n\x00\n\x00'
=== Python 3.13.14, PYTHONUTF8=1 (UTF-8 mode on)
customers OK ['Zoë Müller', 'Renée O’Brien', 'Åsa Lindqvist']
shop FAIL UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9 in position 18: invalid continuation byte
report OK '✓ Zoë Müller (Zürich)\n'
report.txt is 28 bytes after that attempt
native FAIL AttributeError: 'NoneType' object has no attribute 'strip'
ps_export FAIL UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte
[stderr] Exception in thread Thread-1 (_readerthread):
[stderr] UnicodeDecodeError: 'utf-8' codec can't decode byte 0x82 in position 3: invalid start byte
=== Python 3.15.0rc3, default (UTF-8 mode on)
customers OK ['Zoë Müller', 'Renée O’Brien', 'Åsa Lindqvist']
shop FAIL UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9 in position 18: invalid continuation byte
report OK '✓ Zoë Müller (Zürich)\n'
report.txt is 28 bytes after that attempt
native FAIL AttributeError: 'NoneType' object has no attribute 'strip'
ps_export FAIL UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte
[stderr] Exception in thread Thread-1 (_readerthread):
[stderr] UnicodeDecodeError: 'utf-8' codec can't decode byte 0x82 in position 3: invalid start byte
=== Python 3.15.0rc3, PYTHONUTF8=0 (UTF-8 mode off)
customers OK ['Zoë Müller', 'Renée O’Brien', 'Åsa Lindqvist']
shop OK {'name': 'Café Délice', 'owner': 'José Núñez'}
report FAIL UnicodeEncodeError: 'charmap' codec can't encode character '\u2713' in position 0: character maps to <undefined>
report.txt is 0 bytes after that attempt
native OK 'caf‚'
ps_export OK 'ÿþZ\x00o\x00ë\x00 \x00M\x00ü\x00l\x00l\x00e\x00r\x00\n\x00\n\x00'
To keep the outputs in this tutorial short, I removed the folder prefix from file paths and collapsed each traceback to its last line. Your terminal will show the full tracebacks.
Reading the old default (the first and last blocks)
The Python 3.13 default and the Python 3.15 run with PYTHONUTF8=0 print identical results, which confirms again that the mode, not the version, makes the difference. Five lines produce three kinds of outcome:
customersis wrong, silently. The names came back asZoë MüllerandRenée O’Brien. No exception was raised, so a program like this would store garbage in a database and nobody would notice until a customer complained. The curly apostrophe in O’Brien produced the familiar’.reportcrashes withUnicodeEncodeError, because cp1252 has no check mark. Look at the line after it:report.txtis 0 bytes. The callopen(path, "w")creates or empties the file before the first write, so a failed run leaves an empty file where a good report used to be.nativeandps_exportare wrong, silently. The console command wrote the OEM byte82foré, which cp1252 reads as‚. The PowerShell file came back asÿþZ\x00o\x00and so on: the BOM bytes shown as characters, with NUL padding between the letters.shopworks, but only by luck: the file is cp1252 and so is the default.
Whether cp1252 garbles or crashes depends on which letters appear. Create undefined_byte.py:
# undefined_byte.py
data = "Łukasz".encode("utf-8")
print(data.hex(" "))
try:
data.decode("cp1252")
except UnicodeDecodeError as exc:
print(type(exc).__name__, exc)
undefined = []
for value in range(256):
try:
bytes([value]).decode("cp1252")
except UnicodeDecodeError:
undefined.append(f"{value:02x}")
print("bytes with no meaning in cp1252:", " ".join(undefined))
python undefined_byte.py
c5 81 75 6b 61 73 7a
UnicodeDecodeError 'charmap' codec can't decode byte 0x81 in position 1: character maps to <undefined>
bytes with no meaning in cp1252: 81 8d 8f 90 9d
The letter Ł is c5 81 in UTF-8, and byte 81 is one of five that cp1252 leaves undefined, so a UTF-8 file containing that name fails loudly under the old default while a file with ë only garbles. You cannot rely on either behavior.
Reading the new default (the middle two blocks)
With UTF-8 mode on, in the 3.13 run with PYTHONUTF8=1 and in the 3.15 default run, the results are identical to each other and almost the opposite of the first pair:
customersandreportare now correct, and the report file holds 28 bytes including the check mark.shopfails withUnicodeDecodeErroron byte0xe9. The legacy file did not change; the default did. This is the price of the new default: files in the old code page must now be named explicitly.ps_exportfails on byte0xff, the first byte of the UTF-16 BOM. That is an improvement in disguise: the old default returned junk full of NUL characters, and the new one refuses to guess.nativefails withAttributeError: 'NoneType' object has no attribute 'strip', which has nothing obvious to do with encodings. Look at the two[stderr]lines. On Windows,subprocessreads the pipe in a helper thread (the first line names_readerthread). TheUnicodeDecodeErrorfor byte0x82happened inside that thread, sorun()never received it: the thread printed its traceback to stderr andresult.stdoutcame back asNone. When you see an unexplainedNoneafter a subprocess call, read stderr first.
One more place the old default bites: printing
Writing to a pipe or a redirected file with print() also uses the default. Create print_check.py. The matrix captures its output through a pipe, which is how CI logs and scheduled tasks see a program:
# print_check.py
print("✓ report written")
python matrix.py print_check.py
=== Python 3.13.14, default (UTF-8 mode off)
[stderr] UnicodeEncodeError: 'charmap' codec can't encode character '\u2713' in position 0: character maps to <undefined>
[exit 1]
=== Python 3.13.14, PYTHONUTF8=1 (UTF-8 mode on)
✓ report written
=== Python 3.15.0rc3, default (UTF-8 mode on)
✓ report written
=== Python 3.15.0rc3, PYTHONUTF8=0 (UTF-8 mode off)
[stderr] UnicodeEncodeError: 'charmap' codec can't encode character '\u2713' in position 0: character maps to <undefined>
[exit 1]
The first run exits with status 1 and a UnicodeEncodeError. The check mark appears as \u2713 in the message because standard error always uses the backslashreplace handler, which the Python docs for UTF-8 mode confirm: “sys.stderr continues to use backslashreplace as it does in the default locale-aware mode”. My automation has no console window, so every run in this tutorial goes through a pipe, exactly the situation of a redirect such as python tool.py > log.txt.
Step 5: Preview the new default before you upgrade
The second block of the 3.13 run was the 3.15 behavior, produced by setting one environment variable. For your own project you do not need the matrix: run the tests once with UTF-8 mode on and once with it off, and compare. In PowerShell:
$env:PYTHONUTF8 = "1"
python -m pytest -q
Remove-Item Env:PYTHONUTF8
The command-line flag python -X utf8 does the same for a single run, and -X utf8=0 switches UTF-8 mode off on 3.15. There is one trap with the flag. Create inherit_check.py, which starts a second Python and reports its mode:
# inherit_check.py
import subprocess
import sys
print("this process :", sys.flags.utf8_mode)
child = subprocess.run(
[sys.executable, "-c", "import sys; print(sys.flags.utf8_mode)"],
capture_output=True,
text=True,
encoding="ascii",
)
print("child process:", child.stdout.strip())
python -X utf8 inherit_check.py
this process : 1
child process: 0
$env:PYTHONUTF8 = "1"
python inherit_check.py
Remove-Item Env:PYTHONUTF8
this process : 1
child process: 1
In this test on 3.13.14, the flag applies only to the process you started, and the child reports mode 0. The environment variable reaches the child too. If your tests start other Python processes, use the variable. Run these two commands in a shell where PYTHONUTF8 is not already set (the check with Get-ChildItem Env:PYTHON* from Step 2 tells you), or the flag run will report 1 for both processes.
Step 6: Find every implicit default
You now know what the problem looks like. The next question is where it lives in a real code base. No single tool finds everything, so use three.
Tool 1: let Python warn you at runtime
Python can warn whenever the default encoding is used. The io documentation says: “To find where the default encoding is used, you can enable the -X warn_default_encoding command line option or set the PYTHONWARNDEFAULTENCODING environment variable, which will emit an EncodingWarning when the default encoding is used.” Run the original driver with the flag:
python matrix.py -X warn_default_encoding demo.py reportkit_v1
The matrix passes the flag on to each interpreter. Here is the first block, Python 3.13 with the old default:
=== Python 3.13.14, default (UTF-8 mode off)
customers OK ['Zoë Müller', 'Renée O’Brien', 'Åsa Lindqvist']
shop OK {'name': 'Café Délice', 'owner': 'José Núñez'}
report FAIL UnicodeEncodeError: 'charmap' codec can't encode character '\u2713' in position 0: character maps to <undefined>
report.txt is 0 bytes after that attempt
native OK 'caf‚'
ps_export OK 'ÿþZ\x00o\x00ë\x00 \x00M\x00ü\x00l\x00l\x00e\x00r\x00\n\x00\n\x00'
[stderr] reportkit_v1.py:9: EncodingWarning: 'encoding' argument not specified
[stderr] with open(path) as f:
[stderr] reportkit_v1.py:15: EncodingWarning: 'encoding' argument not specified
[stderr] parser.read(path)
[stderr] reportkit_v1.py:20: EncodingWarning: 'encoding' argument not specified
[stderr] with open(path, "w") as f:
[stderr] reportkit_v1.py:30: EncodingWarning: 'encoding' argument not specified.
[stderr] result = subprocess.run(["cmd", "/c", "echo", text], capture_output=True, text=True)
[stderr] reportkit_v1.py:35: EncodingWarning: 'encoding' argument not specified
[stderr] with open(path) as f:
Each warning names the file and line where the encoding was left out, and the next line shows the code. Lines 9, 15, 20, 30 and 35 are reported. Line 26, the read_text() call, is missing, and the reason is instructive: in this run write_report crashed first, so read_report never ran, and a runtime check only sees code that executes. In the blocks where UTF-8 mode is on, line 26 is reported too, and the warnings also appear under the 3.15 default, so the technique keeps working after you upgrade. Line 15, parser.read(path), is reported because configparser opens the file for you, and only a runtime check sees through that.
Tool 2: scan the source code without running it
A static scan reads the code instead of running it, so it finds calls on paths that no test reaches. Python’s standard library includes ast, which turns source text into a tree you can walk. This scanner looks for four patterns, each without an encoding: open() and io.open() in text mode, read_text() and write_text(), and subprocess calls with text=True. Create scan_encoding.py:
# scan_encoding.py
import ast
import sys
from pathlib import Path
SUBPROCESS_FUNCS = {"run", "Popen", "check_output", "check_call", "call"}
def keyword(call, name):
return next((kw for kw in call.keywords if kw.arg == name), None)
def is_true(node):
return isinstance(node, ast.Constant) and node.value is True
def is_binary(call):
mode = None
if len(call.args) > 1 and isinstance(call.args[1], ast.Constant):
mode = call.args[1].value
kw = keyword(call, "mode")
if kw and isinstance(kw.value, ast.Constant):
mode = kw.value.value
return isinstance(mode, str) and "b" in mode
def problem(call):
if any(kw.arg is None for kw in call.keywords) or any(isinstance(a, ast.Starred) for a in call.args):
return None # **kwargs or *args: cannot tell
func = call.func
name = func.id if isinstance(func, ast.Name) else func.attr if isinstance(func, ast.Attribute) else ""
owner = func.value.id if isinstance(func, ast.Attribute) and isinstance(func.value, ast.Name) else ""
if name == "open" and owner in ("", "io"):
if is_binary(call) or keyword(call, "encoding") or len(call.args) >= 4:
return None
return "open() without encoding="
if name in ("read_text", "write_text"):
position = 0 if name == "read_text" else 1
if keyword(call, "encoding") or len(call.args) > position:
return None
return f"{name}() without encoding="
if owner == "subprocess" and name in SUBPROCESS_FUNCS:
text_mode = any(is_true(kw.value) for kw in call.keywords if kw.arg in ("text", "universal_newlines"))
if text_mode and not keyword(call, "encoding"):
return f"subprocess.{name}(text=True) without encoding="
return None
status = 0
for name in sys.argv[1:]:
tree = ast.parse(Path(name).read_text(encoding="utf-8"), filename=name)
findings = sorted(
(node.lineno, message)
for node in ast.walk(tree)
if isinstance(node, ast.Call) and (message := problem(node))
)
for lineno, message in findings:
print(f"{name}:{lineno}: {message}")
print(f"{name}: {len(findings)} finding(s)")
if findings:
status = 1
sys.exit(status)
The function problem() inspects each call node. It skips calls that use *args or **kwargs, because it cannot know what they hold, and it treats a mode containing b as binary. For open() it also accepts a fourth positional argument as an encoding, because open(file, mode, buffering, encoding) puts it there. The script exits with status 1 when it finds anything, so it can serve as a CI gate. Run it on the original:
python scan_encoding.py reportkit_v1.py
reportkit_v1.py:9: open() without encoding=
reportkit_v1.py:20: open() without encoding=
reportkit_v1.py:26: read_text() without encoding=
reportkit_v1.py:30: subprocess.run(text=True) without encoding=
reportkit_v1.py:35: open() without encoding=
reportkit_v1.py: 5 finding(s)
The scanner finds five calls, including the read_text() that the runtime check missed. It misses parser.read(path) on line 15, because a method call on a variable looks like any other method call and the scanner cannot tell that parser is a ConfigParser.
Tool 3: let Ruff check it in CI
Ruff has a rule for this, PLW1514 (unspecified-encoding). It is a preview rule: running ruff rule PLW1514 prints “This rule is in preview and is not stable.” and says the --preview flag is required to use it. Run it on both files (the second file exists after Step 7; the command works as written once you have created it):
ruff check --select PLW1514 --preview --output-format concise reportkit_v1.py reportkit_v2.py
reportkit_v1.py:9:10: unspecified-encoding: `open` in text mode without explicit `encoding` argument
reportkit_v1.py:20:10: unspecified-encoding: `open` in text mode without explicit `encoding` argument
reportkit_v1.py:26:12: unspecified-encoding: `pathlib.Path(...).read_text` without explicit `encoding` argument
reportkit_v1.py:35:10: unspecified-encoding: `open` in text mode without explicit `encoding` argument
Found 4 errors.
No fixes available (4 hidden fixes can be enabled with the `--unsafe-fixes` option).
Ruff reports four of the six sites in the original: lines 9, 20, 26 and 35. It does not report line 15 (parser.read) or line 30 (the subprocess call). Once the rule is in your CI configuration, it flags every new open() that omits the encoding.
Putting the three tools together
| Line | Call | Runtime check, 3.13 old default | Runtime check, UTF-8 mode on | Scanner | Ruff PLW1514 |
|---|---|---|---|---|---|
| 9 | open(path) in load_customers |
yes | yes | yes | yes |
| 15 | parser.read(path) |
yes | yes | no | no |
| 20 | open(path, "w") |
yes | yes | yes | yes |
| 26 | Path(path).read_text() |
no, never ran | yes | yes | yes |
| 30 | subprocess.run(..., text=True) |
yes | yes | yes | no |
| 35 | open(path) in load_ps_export |
yes | yes | yes | yes |
Each tool has a blind spot, and together they cover all six sites. Run the runtime check with your tests, run a static check in CI, and treat a site that only one tool reports as just as real as the rest.
Step 7: Fix each call by naming the encoding of its source
The cure for all six sites is the advice in Python’s own release notes: “This only applies when no encoding argument is given. For best compatibility between versions of Python, ensure that an explicit encoding argument is always provided.” The hard part is choosing the right encoding for each call. The question to ask is not “what is my default?” but “who wrote these bytes?” Create reportkit_v2.py:
# reportkit_v2.py
import configparser
import json
import subprocess
from pathlib import Path
def load_customers(path):
with open(path, encoding="utf-8") as f: # written by a Linux job
return json.load(f)
def load_shop(path):
parser = configparser.ConfigParser()
parser.read(path, encoding="cp1252") # saved by a legacy tool as "ANSI"
return dict(parser["shop"])
def write_report(path, customers):
with open(path, "w", encoding="utf-8") as f: # our own output: UTF-8 on every machine
for c in customers:
f.write(f"✓ {c['name']} ({c['city']})\n")
def read_report(path):
return Path(path).read_text(encoding="utf-8")
def native_echo(text):
# Console programs write in the OEM code page, so decode with "oem".
result = subprocess.run(["cmd", "/c", "echo", text], capture_output=True, text=True, encoding="oem")
return result.stdout.strip()
def load_ps_export(path):
with open(path, encoding="utf-16") as f: # Windows PowerShell 5.1 ">" writes UTF-16 LE with a BOM
return f.read()
Go through the changes one by one:
load_customersnames UTF-8, because a Linux job wrote the file and JSON is a UTF-8 world.load_shoppassesencoding="cp1252"toparser.read. The file is in the Windows code page, and naming that code page gives the same result on every machine. If the file is always produced on the same machine by tools that save in the ANSI code page,encoding="locale"is the alternative; Step 2 showed that it stays cp1252 even in UTF-8 mode, but it follows the machine, so it changes on a computer with different regional settings.write_reportandread_reportuse UTF-8, because your own output should not depend on which machine wrote it.native_echousesencoding="oem". The codecs documentation describes it as “Windows only: Encode the operand according to the OEM codepage (CP_OEMCP).” That matches Microsoft’s statement that OEM code pages are “still used for console applications”. Because the codec exists only on Windows, a program that must also run on Linux needs a platform check around this call.load_ps_exportusesencoding="utf-16". This codec reads the BOM to find the byte order and does not include it in the text. Better still, fix the producer: Microsoft’s advice for controlling output is “When you need to specify parameters for the output, use Out-File rather than the redirection operator”, and Out-File has an-Encodingparameter.
The same decisions in table form, for the next file you write:
| Where the text came from | What to pass | Why |
|---|---|---|
| Files your program writes; JSON, TOML and YAML; anything from Linux, git or the web | encoding="utf-8" |
Same bytes on every machine |
| A legacy file in a known Windows code page | That code page, such as encoding="cp1252" |
Same on every machine |
| A legacy file always produced locally by ANSI tools | encoding="locale" |
Follows the machine; unchanged by UTF-8 mode |
| Output of a console program | encoding="oem" (Windows only) |
Console programs write in the OEM code page |
A file written by Windows PowerShell 5.1 with > |
encoding="utf-16" |
Reads the BOM and drops it |
Run the same driver against the fixed module, and the scanner on it:
python matrix.py demo.py reportkit_v2
python scan_encoding.py reportkit_v2.py
=== Python 3.13.14, default (UTF-8 mode off)
customers OK ['Zoë Müller', 'Renée O’Brien', 'Åsa Lindqvist']
shop OK {'name': 'Café Délice', 'owner': 'José Núñez'}
report OK '✓ Zoë Müller (Zürich)\n'
report.txt is 28 bytes after that attempt
native OK 'café'
ps_export OK 'Zoë Müller'
reportkit_v2.py: 0 finding(s)
All five lines are correct. I show only the first block because the other three blocks of the matrix output are identical to it, which Step 8 turns into a test. The scanner now reports no findings. For a stricter check, make every missing encoding an exception. The flag -W error::EncodingWarning turns the warning into an error, and the driver prints it as a FAIL line:
python matrix.py -X warn_default_encoding -W error::EncodingWarning demo.py reportkit_v2
In all four runs no line prints FAIL and nothing appears on stderr, so the fixed module never relies on a default.
That leaves print(). The driver’s first line, sys.stdout.reconfigure(encoding="utf-8"), is why its output survived the old default in every block. Put the same line at the top of any script that prints non-ASCII text, or set PYTHONUTF8=1 in the environment of the job that runs it. Under UTF-8 mode, the Step 4 print check passed without either.
Step 8: Make the test suite prove it
UTF-8 mode is decided when the interpreter starts, so a test cannot flip it from the inside. The trick is to launch fresh interpreters from the test, one per mode, and compare what they print. Create test_reportkit.py:
# test_reportkit.py
import os
import subprocess
import sys
from pathlib import Path
import pytest
HERE = Path(__file__).parent
NAMES = ["Zoë Müller", "Renée O’Brien", "Åsa Lindqvist"]
@pytest.fixture(scope="session", autouse=True)
def fixtures():
subprocess.run([sys.executable, "make_fixtures.py"], cwd=HERE, check=True, capture_output=True)
def run_demo(module, utf8, strict=False):
env = {k: v for k, v in os.environ.items() if not k.upper().startswith("PYTHON")}
env["PYTHONUTF8"] = utf8
cmd = [sys.executable]
if strict:
cmd += ["-X", "warn_default_encoding", "-W", "error::EncodingWarning"]
proc = subprocess.run(cmd + ["demo.py", module], cwd=HERE, env=env, capture_output=True)
assert proc.returncode == 0, proc.stderr.decode("utf-8", "replace")
return proc.stdout.decode("utf-8")
def run_scan(module):
proc = subprocess.run([sys.executable, "scan_encoding.py", module], cwd=HERE, capture_output=True, text=True, encoding="utf-8")
return proc.returncode
def test_fixed_module_behaves_the_same_in_both_modes():
assert run_demo("reportkit_v2", "0", strict=True) == run_demo("reportkit_v2", "1", strict=True)
def test_fixed_module_reads_every_source_correctly():
out = run_demo("reportkit_v2", "0", strict=True)
assert "FAIL" not in out
assert repr(NAMES) in out
assert "'Café Délice'" in out and "'José Núñez'" in out
assert "'✓ Zoë Müller (Zürich)\\n'" in out
assert "'café'" in out
assert "ps_export OK 'Zoë Müller'" in out
def test_original_module_changes_behaviour_between_modes():
assert run_demo("reportkit_v1", "0") != run_demo("reportkit_v1", "1")
def test_strict_warnings_flag_the_original_module():
out = run_demo("reportkit_v1", "1", strict=True)
assert out.count("EncodingWarning") >= 4
def test_scanner_flags_the_original_but_not_the_fixed_module():
assert run_scan("reportkit_v1.py") == 1
assert run_scan("reportkit_v2.py") == 0
Each test answers one question:
test_fixed_module_behaves_the_same_in_both_modesruns the driver withPYTHONUTF8set to 0 and to 1, with the strict warning flags, and requires identical output. This is the guarantee you want.test_fixed_module_reads_every_source_correctlychecks the actual values, so “identical but identically wrong” cannot pass.test_original_module_changes_behaviour_between_modesruns the original module in both modes and requires different output. It looks odd to test the broken code, but it proves the harness can see mode-dependent behavior at all. If it ever fails, your test setup has gone blind.test_strict_warnings_flag_the_original_modulechecks that the strict flags catch at least four warnings in the original.test_scanner_flags_the_original_but_not_the_fixed_modulechecks the scanner’s exit status, 1 for the original and 0 for the fixed module.
Run the suite under all four interpreter settings:
python matrix.py -m pytest -q test_reportkit.py
=== Python 3.13.14, default (UTF-8 mode off)
..... [100%]
5 passed in 0.94s
=== Python 3.13.14, PYTHONUTF8=1 (UTF-8 mode on)
..... [100%]
5 passed in 0.92s
=== Python 3.15.0rc3, default (UTF-8 mode on)
..... [100%]
5 passed in 0.83s
=== Python 3.15.0rc3, PYTHONUTF8=0 (UTF-8 mode off)
..... [100%]
5 passed in 0.81s
Five tests pass in each of the four runs. The outer interpreter’s mode does not matter, because every test sets the mode of its own child process. For a CI gate that covers your whole suite, run it with the strict flags:
python -X warn_default_encoding -W error::EncodingWarning -m pytest -q test_reportkit.py
..... [100%]
5 passed in 0.94s
That passes here, which means pytest and the standard-library code these tests call raised no EncodingWarning. A larger project may have dependencies that do; the warning points at the file that omitted the encoding, so you can see whether the cause is your code or a library.
Common mistakes and gotchas
- A preset environment hides the bugs.
PYTHONUTF8=1in your shell, your CI image or an IDE run configuration makes 3.13 behave like 3.15, andPYTHONIOENCODING=utf-8changes the standard streams. I had both set without knowing it. Check withGet-ChildItem Env:PYTHON*before you trust any comparison. - Silencing the error is not fixing it. Adding
errors="replace"orerrors="ignore"makes a decode failure disappear by storing�or dropping the character, exactly the loss you saw in Step 1. Name the right encoding instead. - The flag does not reach child Python processes. Use
PYTHONUTF8when tests or tools start other interpreters (Step 5). - A subprocess decode error can surface as
None. On Windows, check stderr whenresult.stdoutisNone(Step 4). - A file that opens fine is not proof of the right encoding. The old default read the UTF-8 customers file without an error and produced wrong names. Compare against known values, as the second test does.
- Each scanner has blind spots. Run the runtime check, a static scan and Ruff together (Step 6).
- Files that start with a BOM need a different codec. A UTF-8 file saved with a BOM, which some Windows tools add, needs
encoding="utf-8-sig"; my CSV parsing tutorial shows that failure and fix.
Check the whole thing end to end
From a clean encoding-lab folder that holds the files from this tutorial, these commands should reproduce the results above:
python make_fixtures.py
python matrix.py demo.py reportkit_v1
python matrix.py demo.py reportkit_v2
python scan_encoding.py reportkit_v1.py
python scan_encoding.py reportkit_v2.py
python matrix.py -m pytest -q test_reportkit.py
You should see four result blocks for the original module that fall into two different pairs, four identical blocks for the fixed one, five scanner findings with exit status 1 and then none with exit status 0, and five passing tests in each of four runs. If the first command prints different byte values, your regional settings differ from mine; the pattern should still hold.
Where to go next
Run scan_encoding.py on your own files, then add Ruff’s PLW1514 rule to CI so new code cannot reintroduce the problem. Set PYTHONWARNDEFAULTENCODING in a CI job that runs your tests, and decide each site with the table from Step 7. Upgrade when you are ready: with every encoding named, the Python version stops mattering.
Three related tutorials on this site pair well with this one. The Python 3.15 lazy imports tutorial and the Tachyon sampling profiler tutorial cover the other two 3.15 features I have tested the same way. Once your bytes decode correctly, the invisible Unicode tutorial shows how to check which characters the text actually contains.








No Comment! Be the first one.