TRENDING
Close-up of the Rosetta Stone showing the Demotic script above and the Greek script below, the same text written in two different scripts
October 6, 2026
How to Prepare Your Python Code for the Python 3.15 UTF-8 Default and Fix Windows Encoding Bugs
A row of green and grey fibre broadband street cabinets on a pavement beside a fence in Iver, England
October 6, 2026
BT’s TalkTalk Rescue Turns Telecom Continuity Into a New Merger-Control Ground
An ornate cast-iron wall mailbox with its door hanging open, stuffed with colorful flyers and a yellow flyer bulging out of the top slot
October 6, 2026
Google Stops Accepting Product Bug Reports for Its Open-Source Bounty, Citing Automated Submissions
Chronophotograph by Étienne-Jules Marey of a man riding a bicycle, showing five snapshots of the same ride taken at regular intervals
October 6, 2026
How to Find Slow Python Code With the Python 3.15 Tachyon Sampling Profiler
Close-up of an airport baggage tag reading Stockholm Arlanda and ARN
October 6, 2026
Cloudflare Traces Turns Distributed Tracing Into a Trust Decision at the Edge
06 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Shelves of old books fastened by iron chains in the Francis Trigge Chained Library in Grantham, England, a picture of data that can be read but not changed
How to Use frozendict in Python 3.15 to Freeze Config and Cache Dictionary Arguments
October 5, 2026
Row of capsule hotel pods with white pillows and folded blankets, each capsule an idle sleeper packed into a shared rack
Kubernetes Node Swap Turns Idle Agent Memory Into a Density Bet With No Wake-Up Test
October 5, 2026
Denmark’s oldest church book, from Holmens parish, open on a stack of books; its handwritten pages record births between 1617 and 1639
Denmark Says 8.8 Million Population Register Records Were Pulled Through One Company’s Lawful Access
October 5, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 226 Posts
News 228 Posts
Learning Hub 198 Posts
Home/Learning Hub/How to Prepare Your Python Code for the Python 3.15 UTF-8 Default and Fix Windows Encoding Bugs
Learning Hub

How to Prepare Your Python Code for the Python 3.15 UTF-8 Default and Fix Windows Encoding Bugs

Python 3.15 makes UTF-8 the default for open(), pipes and stdio. This Windows tutorial uses Python 3.13 and 3.15 RC3 to show which bugs the change fixes, which it exposes, and how to make your code...

October 5, 2026 25 Min Read
9

Python 3.15 changes what open() does when you do not say which text encoding to use. On Windows, that call has fallen back to the machine’s legacy code page, which on my test machine is cp1252. In 3.15 it falls back to UTF-8 instead. The release notes say it directly: “Python now uses UTF-8 as the default encoding, independent of the system’s environment.” The change cures a famous class of Windows bugs, and it also breaks any program that quietly depended on the old default.

Table Of Content

  • The ideas you need first
  • Prerequisites
  • Step 1: See one character as four byte sequences
  • Step 2: Ask Python what it will do by default
  • Step 3: Create test files that come from different sources
  • Step 4: Write the program with implicit defaults and run it under both modes
  • Reading the old default (the first and last blocks)
  • Reading the new default (the middle two blocks)
  • One more place the old default bites: printing
  • Step 5: Preview the new default before you upgrade
  • Step 6: Find every implicit default
  • Tool 1: let Python warn you at runtime
  • Tool 2: scan the source code without running it
  • Tool 3: let Ruff check it in CI
  • Putting the three tools together
  • Step 7: Fix each call by naming the encoding of its source
  • Step 8: Make the test suite prove it
  • Common mistakes and gotchas
  • Check the whole thing end to end
  • Where to go next

In this tutorial you will build a small report tool that reads and writes text from four different sources, run it under the old default and the new default on a real Windows 11 machine, and watch one setting turn the same code into correct output, garbled text, or a crash. Then you will find every implicit default with three different tools, fix each call by naming the encoding of its source, and add a test suite that proves the fixed code behaves identically under both defaults. You can start today without installing Python 3.15, because one environment variable previews the new behavior on any recent Python.

The ideas you need first

A file on disk holds bytes, not letters. An encoding is the rulebook that maps letters to bytes and back. Reading bytes with the wrong rulebook produces mojibake, the garbled text you get when, for example, é turns into é. When your code calls open("notes.txt") without an encoding= argument, Python has to guess the rulebook. That guess is the default encoding.

Before 3.15, the guess on Windows came from the locale encoding, the code page configured for the system. According to Microsoft’s documentation, Windows code pages are commonly called ANSI code pages, and code page 1252 is commonly used for English and other Western European languages. Console programs use a different one, the OEM code page. The same page says of OEM code pages: “These code pages were originally used for MS-DOS and are still used for console applications.” It adds: “The usual OEM code page for English is code page 437.” So one Windows machine can hold text in four encodings at once: UTF-8 from Linux, git and the web; ANSI from older tools; OEM from console programs; and UTF-16 from some Microsoft tools.

Python 3.15 turns on UTF-8 mode by default. PEP 686 states the goal: “With this change, Python consistently uses UTF-8 for default encoding of files, stdio, and pipes.” Python source files were already read as UTF-8 (“The default encoding of Python source files is UTF-8”, says the same PEP), so the change brings your data files in line with your source files.

Prerequisites

  • Windows 10 or 11 and a PowerShell terminal. I used Windows 11 (version 10.0.26100), where the ANSI code page is cp1252 and the OEM code page is 437, the usual setup for English. Other regional settings use other code pages, so your exact bytes may differ, but the lessons are the same. Step 2 shows how to read your own settings.
  • Python 3.11 or newer. The function locale.getencoding(), which this tutorial uses, was added in 3.11. I used 3.13.14.
  • Optional: a Python 3.15 build. I used 3.15.0rc3, the newest build in python.org’s 3.15.0 download folder on October 5, 2026. PEP 790 lists the final release as expected on Friday, 2026-10-09. You do not need an installer: the zip file runs in place.
  • pytest and Ruff, installed with pip. I used pytest 9.1.1 and Ruff 0.16.10.
  • Basic Python: functions, files, and reading a traceback. Everything else is explained as it appears.

Set up a working folder with a virtual environment and the optional 3.15 build. Putting the virtual environment on the search path for the current session avoids the PowerShell script-execution policy that blocks Activate.ps1 on many machines:

mkdir encoding-lab
cd encoding-lab
python -m venv .venv
$env:PATH = "$PWD\.venv\Scripts;$env:PATH"
python -m pip install pytest ruff
curl.exe -L -o py315.zip https://www.python.org/ftp/python/3.15.0/python-3.15.0rc3-amd64.zip
Expand-Archive py315.zip -DestinationPath py315
.\py315\python.exe -m ensurepip --upgrade
.\py315\python.exe -m pip install pytest
.\py315\python.exe --version

The last command should print Python 3.15.0rc3. Recent versions of Windows include curl.exe; type the .exe, because in Windows PowerShell 5.1 plain curl is an alias for a different command. Every file below goes into encoding-lab, and the first line of each code block (a comment such as # matrix.py) is the file name.

Step 1: See one character as four byte sequences

The quickest way to understand every bug in this tutorial is to look at one character, é, in four encodings. Create same_char.py:

# same_char.py
import sys

sys.stdout.reconfigure(encoding="utf-8")  # this demo prints odd characters, so pin its own output

text = "é"
print(f"{'encoding':<10} {'bytes':<8} {'read as cp1252':<16} read as UTF-8")
for name in ("utf-8", "cp1252", "cp437", "utf-16-le"):
    data = text.encode(name)
    as_cp1252 = data.decode("cp1252", errors="replace")
    as_utf8 = data.decode("utf-8", errors="replace")
    print(f"{name:<10} {data.hex(' '):<8} {as_cp1252!r:<16} {as_utf8!r}")

The script encodes é four ways, then decodes each result with the wrong rulebooks on purpose (errors="replace" makes Python substitute the replacement character � instead of raising an error). The line after the imports pins the script’s own output to UTF-8, so it can print any character on any Python. Run it:

python same_char.py
encoding   bytes    read as cp1252   read as UTF-8
utf-8      c3 a9    'é'             'é'
cp1252     e9       'é'              '�'
cp437      82       '‚'              '�'
utf-16-le  e9 00    'é\x00'          '�\x00'

Read the rows one at a time. UTF-8 stores é as the two bytes c3 a9, and a program that reads those bytes as cp1252 sees two characters, Ã and ©. That is the classic mojibake. The cp1252 row is the reverse: one byte, e9, which UTF-8 cannot accept on its own, so a strict UTF-8 reader refuses it. The cp437 row is what a console program writes for é (byte 82); read as cp1252 it becomes the low quotation mark ‚. The last row is UTF-16, where each common character takes two bytes and an ASCII letter carries a zero byte beside it.

Keep this table in mind. Every bug below is one of these rows read through the wrong column.

Step 2: Ask Python what it will do by default

Create settings.py. It prints the encoding settings that matter, including what open() uses when you give it no encoding= argument:

# settings.py
import locale
import os
import sys

with open(os.devnull) as devnull:
    open_default = devnull.encoding
with open(os.devnull, encoding="locale") as devnull:
    open_locale = devnull.encoding

rows = [
    ("Python", sys.version.split()[0]),
    ("locale.getencoding()", locale.getencoding()),
    ("locale.getpreferredencoding(False)", locale.getpreferredencoding(False)),
    ("sys.flags.utf8_mode", sys.flags.utf8_mode),
    ("open(...).encoding, no encoding=", open_default),
    ('open(..., encoding="locale").encoding', open_locale),
    ("sys.getfilesystemencoding()", sys.getfilesystemencoding()),
    ("sys.getdefaultencoding()", sys.getdefaultencoding()),
]
width = max(len(name) for name, _ in rows)
for name, value in rows:
    print(f"{name:<{width}}  {value}")

You also need a way to run any script under both defaults, and under both Python versions, without your shell environment interfering. Create matrix.py:

# matrix.py
import os
import subprocess
import sys
from pathlib import Path

sys.stdout.reconfigure(encoding="utf-8")

args = sys.argv[1:]
pythons = []
while args and args[0] == "--python":
    pythons.append(args[1])
    args = args[2:]
if not pythons:
    pythons = [sys.executable]
    candidate = Path(__file__).parent / "py315" / "python.exe"
    if candidate.exists():
        pythons.append(str(candidate))


def clean_env(utf8=None):
    env = {k: v for k, v in os.environ.items() if not k.upper().startswith("PYTHON")}
    if utf8 is not None:
        env["PYTHONUTF8"] = utf8
    return env


def ask(exe, code):
    out = subprocess.run([exe, "-c", code], env=clean_env(), capture_output=True)
    return out.stdout.decode("ascii", "replace").strip()


for exe in pythons:
    version = ask(exe, "import sys; print(sys.version.split()[0])")
    default_on = ask(exe, "import sys; print(sys.flags.utf8_mode)") == "1"
    flip = "0" if default_on else "1"
    for label, utf8, mode_on in [("default", None, default_on), (f"PYTHONUTF8={flip}", flip, not default_on)]:
        proc = subprocess.run([exe, *args], env=clean_env(utf8), capture_output=True)
        print(f"=== Python {version}, {label} (UTF-8 mode {'on' if mode_on else 'off'})")
        out = proc.stdout.decode("utf-8", "backslashreplace").replace("\r\n", "\n").rstrip()
        if out:
            print(out)
        for line in proc.stderr.decode("utf-8", "backslashreplace").strip().splitlines():
            print("[stderr]", line)
        if proc.returncode:
            print(f"[exit {proc.returncode}]")
        print()

This helper does four things. It finds the interpreter that runs it, plus a Python 3.15 interpreter in py315\python.exe if that folder exists. For each interpreter it asks, with a clean environment, whether UTF-8 mode is on by default. It then runs your command twice: once with the interpreter’s own default, and once with PYTHONUTF8 set to the opposite value. And it first removes every environment variable that starts with PYTHON from the child process, because those variables change the result. While preparing this tutorial I found PYTHONUTF8=1 and PYTHONIOENCODING=utf-8 already set in my own shell. With PYTHONUTF8=1 set, Python 3.13 already behaves like 3.15, so the old-default bugs below never show up. Check yours with Get-ChildItem Env:PYTHON*.

python matrix.py settings.py

The output has one block per run. Here it is reorganized as a table, one column per run:

Setting 3.13.14 default 3.13.14 with PYTHONUTF8=1 3.15.0rc3 default 3.15.0rc3 with PYTHONUTF8=0
sys.flags.utf8_mode 0 1 1 0
locale.getencoding() cp1252 cp1252 cp1252 cp1252
locale.getpreferredencoding(False) cp1252 utf-8 utf-8 cp1252
open(...).encoding, no encoding= cp1252 utf-8 utf-8 cp1252
open(..., encoding="locale").encoding cp1252 cp1252 cp1252 cp1252
sys.getfilesystemencoding() utf-8 utf-8 utf-8 utf-8
sys.getdefaultencoding() utf-8 utf-8 utf-8 utf-8

Read it by columns. The UTF-8 mode row decides what open() does, and the Python version does not matter: 3.13 with PYTHONUTF8=1 matches the 3.15 default, and 3.15 with PYTHONUTF8=0 matches the 3.13 default. The Python docs confirm both halves: “Changed in version 3.15: Python UTF-8 mode is now enabled by default (PEP 686). It may be disabled by setting PYTHONUTF8=0 as an environment variable or by using the -X utf8=0 command line option.”

Three rows deserve a closer look. First, locale.getencoding() says cp1252 in all four runs because, in the docs’ words, “This function is similar to getpreferredencoding(False) except this function ignores the Python UTF-8 Mode.” It tells you what the machine is configured for, whatever Python decides to do. Second, open(..., encoding="locale") still gives cp1252 under UTF-8 mode, which makes it the explicit way to ask for the old behavior. Third, sys.getdefaultencoding() says utf-8 everywhere, even on the old default. It describes how Python converts between str and bytes in memory, not how it reads files, so it is not a useful check here.

If locale.getencoding() on your machine already says utf-8, the two defaults agree and the old-default bugs below will not reproduce for you. Microsoft documents that the ANSI code page can be configured for UTF-8 in its guide to the UTF-8 code page.

Step 3: Create test files that come from different sources

Real programs read files written by other programs. The next script creates three small files that imitate three sources you will meet on Windows. Notice that this script passes an explicit encoding every time it writes text; that is the habit this tutorial builds. Create make_fixtures.py:

# make_fixtures.py
import json
import subprocess
from pathlib import Path

customers = [
    {"name": "Zoë Müller", "city": "Zürich"},
    {"name": "Renée O’Brien", "city": "Cork"},
    {"name": "Åsa Lindqvist", "city": "Malmö"},
]

# 1. What a Linux CI job or a web API writes: UTF-8.
Path("customers.json").write_text(json.dumps(customers, ensure_ascii=False, indent=2), encoding="utf-8")

# 2. What a legacy Windows tool writes when told to save as "ANSI": cp1252 on this machine.
Path("shop.ini").write_text("[shop]\nname = Café Délice\nowner = José Núñez\n", encoding="cp1252")

# 3. What Windows PowerShell 5.1 writes for the > operator: UTF-16 LE with a byte order mark.
subprocess.run(["powershell", "-NoProfile", "-Command", '"Zoë Müller" > ps_export.txt'], check=True)

for name in ("customers.json", "shop.ini", "ps_export.txt"):
    data = Path(name).read_bytes()
    offset = next(i for i, b in enumerate(data) if b >= 0x80)
    print(f"{name:<15} {len(data):>3} bytes, first non-ASCII byte at {offset:>2}: {data[offset:offset + 6].hex(' ')}")
python make_fixtures.py
customers.json  194 bytes, first non-ASCII byte at 23: c3 ab 20 4d c3 bc
shop.ini         48 bytes, first non-ASCII byte at 18: e9 20 44 e9 6c 69
ps_export.txt    26 bytes, first non-ASCII byte at  0: ff fe 5a 00 6f 00

Each line prints the first byte above 127, the first byte that is not plain ASCII, and the bytes after it. In customers.json it is c3 ab, the UTF-8 form of ë. In shop.ini it is a lone e9, the cp1252 form of é. In ps_export.txt it is ff fe at offset 0, a byte order mark (BOM) that announces UTF-16 little-endian, followed by 5a 00 6f 00, the letters Z and o, each padded with a zero byte.

The third file comes from powershell, which on my machine is Windows PowerShell 5.1. Microsoft’s documentation says that redirecting with the > operator is “functionally equivalent to piping to Out-File with no extra parameters”, and the bytes above show what that cmdlet writes by default: UTF-16.

Step 4: Write the program with implicit defaults and run it under both modes

Here is a small report tool written the way most Python code is written, with no encoding= anywhere. Create reportkit_v1.py:

# reportkit_v1.py
import configparser
import json
import subprocess
from pathlib import Path


def load_customers(path):
    with open(path) as f:
        return json.load(f)


def load_shop(path):
    parser = configparser.ConfigParser()
    parser.read(path)
    return dict(parser["shop"])


def write_report(path, customers):
    with open(path, "w") as f:
        for c in customers:
            f.write(f"✓ {c['name']} ({c['city']})\n")


def read_report(path):
    return Path(path).read_text()


def native_echo(text):
    result = subprocess.run(["cmd", "/c", "echo", text], capture_output=True, text=True)
    return result.stdout.strip()


def load_ps_export(path):
    with open(path) as f:
        return f.read()

It has five jobs. load_customers reads JSON that a Linux job produced. load_shop reads a legacy settings file with configparser. write_report and read_report write a report that uses a check mark and read it back. native_echo runs a console command and captures its output; the subprocess docs say the pipes are “opened in text mode using the specified encoding and errors or the io.TextIOWrapper default”, which means the default encoding again. load_ps_export reads the PowerShell file.

Now a driver that calls each function inside try/except, so one failure does not hide the others. Create demo.py:

# demo.py
import importlib
import sys
from pathlib import Path

sys.stdout.reconfigure(encoding="utf-8")  # let the demo print any character, whatever the default is

kit = importlib.import_module(sys.argv[1])
report = Path("report.txt")


def attempt(label, func):
    try:
        result = func()
    except Exception as exc:
        print(f"{label:<10} FAIL  {type(exc).__name__}: {exc}")
    else:
        print(f"{label:<10} OK    {result!r}")


def roundtrip():
    report.unlink(missing_ok=True)
    kit.write_report(report, [{"name": "Zoë Müller", "city": "Zürich"}])
    return kit.read_report(report)


attempt("customers", lambda: [c["name"] for c in kit.load_customers("customers.json")])
attempt("shop", lambda: kit.load_shop("shop.ini"))
attempt("report", roundtrip)
print(f"report.txt is {report.stat().st_size} bytes after that attempt")
attempt("native", lambda: kit.native_echo("café"))
attempt("ps_export", lambda: kit.load_ps_export("ps_export.txt").strip())

The module name is an argument, so you can point the same driver at the fixed version later. Run it under all four combinations:

python matrix.py demo.py reportkit_v1
=== Python 3.13.14, default (UTF-8 mode off)
customers  OK    ['Zoë Müller', 'Renée O’Brien', 'Åsa Lindqvist']
shop       OK    {'name': 'Café Délice', 'owner': 'José Núñez'}
report     FAIL  UnicodeEncodeError: 'charmap' codec can't encode character '\u2713' in position 0: character maps to <undefined>
report.txt is 0 bytes after that attempt
native     OK    'caf‚'
ps_export  OK    'ÿþZ\x00o\x00ë\x00 \x00M\x00ü\x00l\x00l\x00e\x00r\x00\n\x00\n\x00'

=== Python 3.13.14, PYTHONUTF8=1 (UTF-8 mode on)
customers  OK    ['Zoë Müller', 'Renée O’Brien', 'Åsa Lindqvist']
shop       FAIL  UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9 in position 18: invalid continuation byte
report     OK    '✓ Zoë Müller (Zürich)\n'
report.txt is 28 bytes after that attempt
native     FAIL  AttributeError: 'NoneType' object has no attribute 'strip'
ps_export  FAIL  UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte
[stderr] Exception in thread Thread-1 (_readerthread):
[stderr] UnicodeDecodeError: 'utf-8' codec can't decode byte 0x82 in position 3: invalid start byte

=== Python 3.15.0rc3, default (UTF-8 mode on)
customers  OK    ['Zoë Müller', 'Renée O’Brien', 'Åsa Lindqvist']
shop       FAIL  UnicodeDecodeError: 'utf-8' codec can't decode byte 0xe9 in position 18: invalid continuation byte
report     OK    '✓ Zoë Müller (Zürich)\n'
report.txt is 28 bytes after that attempt
native     FAIL  AttributeError: 'NoneType' object has no attribute 'strip'
ps_export  FAIL  UnicodeDecodeError: 'utf-8' codec can't decode byte 0xff in position 0: invalid start byte
[stderr] Exception in thread Thread-1 (_readerthread):
[stderr] UnicodeDecodeError: 'utf-8' codec can't decode byte 0x82 in position 3: invalid start byte

=== Python 3.15.0rc3, PYTHONUTF8=0 (UTF-8 mode off)
customers  OK    ['Zoë Müller', 'Renée O’Brien', 'Åsa Lindqvist']
shop       OK    {'name': 'Café Délice', 'owner': 'José Núñez'}
report     FAIL  UnicodeEncodeError: 'charmap' codec can't encode character '\u2713' in position 0: character maps to <undefined>
report.txt is 0 bytes after that attempt
native     OK    'caf‚'
ps_export  OK    'ÿþZ\x00o\x00ë\x00 \x00M\x00ü\x00l\x00l\x00e\x00r\x00\n\x00\n\x00'

To keep the outputs in this tutorial short, I removed the folder prefix from file paths and collapsed each traceback to its last line. Your terminal will show the full tracebacks.

Reading the old default (the first and last blocks)

The Python 3.13 default and the Python 3.15 run with PYTHONUTF8=0 print identical results, which confirms again that the mode, not the version, makes the difference. Five lines produce three kinds of outcome:

  • customers is wrong, silently. The names came back as Zoë Müller and Renée O’Brien. No exception was raised, so a program like this would store garbage in a database and nobody would notice until a customer complained. The curly apostrophe in O’Brien produced the familiar ’.
  • report crashes with UnicodeEncodeError, because cp1252 has no check mark. Look at the line after it: report.txt is 0 bytes. The call open(path, "w") creates or empties the file before the first write, so a failed run leaves an empty file where a good report used to be.
  • native and ps_export are wrong, silently. The console command wrote the OEM byte 82 for é, which cp1252 reads as ‚. The PowerShell file came back as ÿþZ\x00o\x00 and so on: the BOM bytes shown as characters, with NUL padding between the letters.
  • shop works, but only by luck: the file is cp1252 and so is the default.

Whether cp1252 garbles or crashes depends on which letters appear. Create undefined_byte.py:

# undefined_byte.py
data = "Łukasz".encode("utf-8")
print(data.hex(" "))
try:
    data.decode("cp1252")
except UnicodeDecodeError as exc:
    print(type(exc).__name__, exc)

undefined = []
for value in range(256):
    try:
        bytes([value]).decode("cp1252")
    except UnicodeDecodeError:
        undefined.append(f"{value:02x}")
print("bytes with no meaning in cp1252:", " ".join(undefined))
python undefined_byte.py
c5 81 75 6b 61 73 7a
UnicodeDecodeError 'charmap' codec can't decode byte 0x81 in position 1: character maps to <undefined>
bytes with no meaning in cp1252: 81 8d 8f 90 9d

The letter Ł is c5 81 in UTF-8, and byte 81 is one of five that cp1252 leaves undefined, so a UTF-8 file containing that name fails loudly under the old default while a file with ë only garbles. You cannot rely on either behavior.

Reading the new default (the middle two blocks)

With UTF-8 mode on, in the 3.13 run with PYTHONUTF8=1 and in the 3.15 default run, the results are identical to each other and almost the opposite of the first pair:

  • customers and report are now correct, and the report file holds 28 bytes including the check mark.
  • shop fails with UnicodeDecodeError on byte 0xe9. The legacy file did not change; the default did. This is the price of the new default: files in the old code page must now be named explicitly.
  • ps_export fails on byte 0xff, the first byte of the UTF-16 BOM. That is an improvement in disguise: the old default returned junk full of NUL characters, and the new one refuses to guess.
  • native fails with AttributeError: 'NoneType' object has no attribute 'strip', which has nothing obvious to do with encodings. Look at the two [stderr] lines. On Windows, subprocess reads the pipe in a helper thread (the first line names _readerthread). The UnicodeDecodeError for byte 0x82 happened inside that thread, so run() never received it: the thread printed its traceback to stderr and result.stdout came back as None. When you see an unexplained None after a subprocess call, read stderr first.

One more place the old default bites: printing

Writing to a pipe or a redirected file with print() also uses the default. Create print_check.py. The matrix captures its output through a pipe, which is how CI logs and scheduled tasks see a program:

# print_check.py
print("✓ report written")
python matrix.py print_check.py
=== Python 3.13.14, default (UTF-8 mode off)
[stderr] UnicodeEncodeError: 'charmap' codec can't encode character '\u2713' in position 0: character maps to <undefined>
[exit 1]

=== Python 3.13.14, PYTHONUTF8=1 (UTF-8 mode on)
✓ report written

=== Python 3.15.0rc3, default (UTF-8 mode on)
✓ report written

=== Python 3.15.0rc3, PYTHONUTF8=0 (UTF-8 mode off)
[stderr] UnicodeEncodeError: 'charmap' codec can't encode character '\u2713' in position 0: character maps to <undefined>
[exit 1]

The first run exits with status 1 and a UnicodeEncodeError. The check mark appears as \u2713 in the message because standard error always uses the backslashreplace handler, which the Python docs for UTF-8 mode confirm: “sys.stderr continues to use backslashreplace as it does in the default locale-aware mode”. My automation has no console window, so every run in this tutorial goes through a pipe, exactly the situation of a redirect such as python tool.py > log.txt.

Step 5: Preview the new default before you upgrade

The second block of the 3.13 run was the 3.15 behavior, produced by setting one environment variable. For your own project you do not need the matrix: run the tests once with UTF-8 mode on and once with it off, and compare. In PowerShell:

$env:PYTHONUTF8 = "1"
python -m pytest -q
Remove-Item Env:PYTHONUTF8

The command-line flag python -X utf8 does the same for a single run, and -X utf8=0 switches UTF-8 mode off on 3.15. There is one trap with the flag. Create inherit_check.py, which starts a second Python and reports its mode:

# inherit_check.py
import subprocess
import sys

print("this process :", sys.flags.utf8_mode)
child = subprocess.run(
    [sys.executable, "-c", "import sys; print(sys.flags.utf8_mode)"],
    capture_output=True,
    text=True,
    encoding="ascii",
)
print("child process:", child.stdout.strip())
python -X utf8 inherit_check.py
this process : 1
child process: 0
$env:PYTHONUTF8 = "1"
python inherit_check.py
Remove-Item Env:PYTHONUTF8
this process : 1
child process: 1

In this test on 3.13.14, the flag applies only to the process you started, and the child reports mode 0. The environment variable reaches the child too. If your tests start other Python processes, use the variable. Run these two commands in a shell where PYTHONUTF8 is not already set (the check with Get-ChildItem Env:PYTHON* from Step 2 tells you), or the flag run will report 1 for both processes.

Step 6: Find every implicit default

You now know what the problem looks like. The next question is where it lives in a real code base. No single tool finds everything, so use three.

Tool 1: let Python warn you at runtime

Python can warn whenever the default encoding is used. The io documentation says: “To find where the default encoding is used, you can enable the -X warn_default_encoding command line option or set the PYTHONWARNDEFAULTENCODING environment variable, which will emit an EncodingWarning when the default encoding is used.” Run the original driver with the flag:

python matrix.py -X warn_default_encoding demo.py reportkit_v1

The matrix passes the flag on to each interpreter. Here is the first block, Python 3.13 with the old default:

=== Python 3.13.14, default (UTF-8 mode off)
customers  OK    ['Zoë Müller', 'Renée O’Brien', 'Åsa Lindqvist']
shop       OK    {'name': 'Café Délice', 'owner': 'José Núñez'}
report     FAIL  UnicodeEncodeError: 'charmap' codec can't encode character '\u2713' in position 0: character maps to <undefined>
report.txt is 0 bytes after that attempt
native     OK    'caf‚'
ps_export  OK    'ÿþZ\x00o\x00ë\x00 \x00M\x00ü\x00l\x00l\x00e\x00r\x00\n\x00\n\x00'
[stderr] reportkit_v1.py:9: EncodingWarning: 'encoding' argument not specified
[stderr]   with open(path) as f:
[stderr] reportkit_v1.py:15: EncodingWarning: 'encoding' argument not specified
[stderr]   parser.read(path)
[stderr] reportkit_v1.py:20: EncodingWarning: 'encoding' argument not specified
[stderr]   with open(path, "w") as f:
[stderr] reportkit_v1.py:30: EncodingWarning: 'encoding' argument not specified.
[stderr]   result = subprocess.run(["cmd", "/c", "echo", text], capture_output=True, text=True)
[stderr] reportkit_v1.py:35: EncodingWarning: 'encoding' argument not specified
[stderr]   with open(path) as f:

Each warning names the file and line where the encoding was left out, and the next line shows the code. Lines 9, 15, 20, 30 and 35 are reported. Line 26, the read_text() call, is missing, and the reason is instructive: in this run write_report crashed first, so read_report never ran, and a runtime check only sees code that executes. In the blocks where UTF-8 mode is on, line 26 is reported too, and the warnings also appear under the 3.15 default, so the technique keeps working after you upgrade. Line 15, parser.read(path), is reported because configparser opens the file for you, and only a runtime check sees through that.

Tool 2: scan the source code without running it

A static scan reads the code instead of running it, so it finds calls on paths that no test reaches. Python’s standard library includes ast, which turns source text into a tree you can walk. This scanner looks for four patterns, each without an encoding: open() and io.open() in text mode, read_text() and write_text(), and subprocess calls with text=True. Create scan_encoding.py:

# scan_encoding.py
import ast
import sys
from pathlib import Path

SUBPROCESS_FUNCS = {"run", "Popen", "check_output", "check_call", "call"}


def keyword(call, name):
    return next((kw for kw in call.keywords if kw.arg == name), None)


def is_true(node):
    return isinstance(node, ast.Constant) and node.value is True


def is_binary(call):
    mode = None
    if len(call.args) > 1 and isinstance(call.args[1], ast.Constant):
        mode = call.args[1].value
    kw = keyword(call, "mode")
    if kw and isinstance(kw.value, ast.Constant):
        mode = kw.value.value
    return isinstance(mode, str) and "b" in mode


def problem(call):
    if any(kw.arg is None for kw in call.keywords) or any(isinstance(a, ast.Starred) for a in call.args):
        return None  # **kwargs or *args: cannot tell
    func = call.func
    name = func.id if isinstance(func, ast.Name) else func.attr if isinstance(func, ast.Attribute) else ""
    owner = func.value.id if isinstance(func, ast.Attribute) and isinstance(func.value, ast.Name) else ""
    if name == "open" and owner in ("", "io"):
        if is_binary(call) or keyword(call, "encoding") or len(call.args) >= 4:
            return None
        return "open() without encoding="
    if name in ("read_text", "write_text"):
        position = 0 if name == "read_text" else 1
        if keyword(call, "encoding") or len(call.args) > position:
            return None
        return f"{name}() without encoding="
    if owner == "subprocess" and name in SUBPROCESS_FUNCS:
        text_mode = any(is_true(kw.value) for kw in call.keywords if kw.arg in ("text", "universal_newlines"))
        if text_mode and not keyword(call, "encoding"):
            return f"subprocess.{name}(text=True) without encoding="
    return None


status = 0
for name in sys.argv[1:]:
    tree = ast.parse(Path(name).read_text(encoding="utf-8"), filename=name)
    findings = sorted(
        (node.lineno, message)
        for node in ast.walk(tree)
        if isinstance(node, ast.Call) and (message := problem(node))
    )
    for lineno, message in findings:
        print(f"{name}:{lineno}: {message}")
    print(f"{name}: {len(findings)} finding(s)")
    if findings:
        status = 1
sys.exit(status)

The function problem() inspects each call node. It skips calls that use *args or **kwargs, because it cannot know what they hold, and it treats a mode containing b as binary. For open() it also accepts a fourth positional argument as an encoding, because open(file, mode, buffering, encoding) puts it there. The script exits with status 1 when it finds anything, so it can serve as a CI gate. Run it on the original:

python scan_encoding.py reportkit_v1.py
reportkit_v1.py:9: open() without encoding=
reportkit_v1.py:20: open() without encoding=
reportkit_v1.py:26: read_text() without encoding=
reportkit_v1.py:30: subprocess.run(text=True) without encoding=
reportkit_v1.py:35: open() without encoding=
reportkit_v1.py: 5 finding(s)

The scanner finds five calls, including the read_text() that the runtime check missed. It misses parser.read(path) on line 15, because a method call on a variable looks like any other method call and the scanner cannot tell that parser is a ConfigParser.

Tool 3: let Ruff check it in CI

Ruff has a rule for this, PLW1514 (unspecified-encoding). It is a preview rule: running ruff rule PLW1514 prints “This rule is in preview and is not stable.” and says the --preview flag is required to use it. Run it on both files (the second file exists after Step 7; the command works as written once you have created it):

ruff check --select PLW1514 --preview --output-format concise reportkit_v1.py reportkit_v2.py
reportkit_v1.py:9:10: unspecified-encoding: `open` in text mode without explicit `encoding` argument
reportkit_v1.py:20:10: unspecified-encoding: `open` in text mode without explicit `encoding` argument
reportkit_v1.py:26:12: unspecified-encoding: `pathlib.Path(...).read_text` without explicit `encoding` argument
reportkit_v1.py:35:10: unspecified-encoding: `open` in text mode without explicit `encoding` argument
Found 4 errors.
No fixes available (4 hidden fixes can be enabled with the `--unsafe-fixes` option).

Ruff reports four of the six sites in the original: lines 9, 20, 26 and 35. It does not report line 15 (parser.read) or line 30 (the subprocess call). Once the rule is in your CI configuration, it flags every new open() that omits the encoding.

Putting the three tools together

Line Call Runtime check, 3.13 old default Runtime check, UTF-8 mode on Scanner Ruff PLW1514
9 open(path) in load_customers yes yes yes yes
15 parser.read(path) yes yes no no
20 open(path, "w") yes yes yes yes
26 Path(path).read_text() no, never ran yes yes yes
30 subprocess.run(..., text=True) yes yes yes no
35 open(path) in load_ps_export yes yes yes yes

Each tool has a blind spot, and together they cover all six sites. Run the runtime check with your tests, run a static check in CI, and treat a site that only one tool reports as just as real as the rest.

Step 7: Fix each call by naming the encoding of its source

The cure for all six sites is the advice in Python’s own release notes: “This only applies when no encoding argument is given. For best compatibility between versions of Python, ensure that an explicit encoding argument is always provided.” The hard part is choosing the right encoding for each call. The question to ask is not “what is my default?” but “who wrote these bytes?” Create reportkit_v2.py:

# reportkit_v2.py
import configparser
import json
import subprocess
from pathlib import Path


def load_customers(path):
    with open(path, encoding="utf-8") as f:  # written by a Linux job
        return json.load(f)


def load_shop(path):
    parser = configparser.ConfigParser()
    parser.read(path, encoding="cp1252")  # saved by a legacy tool as "ANSI"
    return dict(parser["shop"])


def write_report(path, customers):
    with open(path, "w", encoding="utf-8") as f:  # our own output: UTF-8 on every machine
        for c in customers:
            f.write(f"✓ {c['name']} ({c['city']})\n")


def read_report(path):
    return Path(path).read_text(encoding="utf-8")


def native_echo(text):
    # Console programs write in the OEM code page, so decode with "oem".
    result = subprocess.run(["cmd", "/c", "echo", text], capture_output=True, text=True, encoding="oem")
    return result.stdout.strip()


def load_ps_export(path):
    with open(path, encoding="utf-16") as f:  # Windows PowerShell 5.1 ">" writes UTF-16 LE with a BOM
        return f.read()

Go through the changes one by one:

  • load_customers names UTF-8, because a Linux job wrote the file and JSON is a UTF-8 world.
  • load_shop passes encoding="cp1252" to parser.read. The file is in the Windows code page, and naming that code page gives the same result on every machine. If the file is always produced on the same machine by tools that save in the ANSI code page, encoding="locale" is the alternative; Step 2 showed that it stays cp1252 even in UTF-8 mode, but it follows the machine, so it changes on a computer with different regional settings.
  • write_report and read_report use UTF-8, because your own output should not depend on which machine wrote it.
  • native_echo uses encoding="oem". The codecs documentation describes it as “Windows only: Encode the operand according to the OEM codepage (CP_OEMCP).” That matches Microsoft’s statement that OEM code pages are “still used for console applications”. Because the codec exists only on Windows, a program that must also run on Linux needs a platform check around this call.
  • load_ps_export uses encoding="utf-16". This codec reads the BOM to find the byte order and does not include it in the text. Better still, fix the producer: Microsoft’s advice for controlling output is “When you need to specify parameters for the output, use Out-File rather than the redirection operator”, and Out-File has an -Encoding parameter.

The same decisions in table form, for the next file you write:

Where the text came from What to pass Why
Files your program writes; JSON, TOML and YAML; anything from Linux, git or the web encoding="utf-8" Same bytes on every machine
A legacy file in a known Windows code page That code page, such as encoding="cp1252" Same on every machine
A legacy file always produced locally by ANSI tools encoding="locale" Follows the machine; unchanged by UTF-8 mode
Output of a console program encoding="oem" (Windows only) Console programs write in the OEM code page
A file written by Windows PowerShell 5.1 with > encoding="utf-16" Reads the BOM and drops it

Run the same driver against the fixed module, and the scanner on it:

python matrix.py demo.py reportkit_v2
python scan_encoding.py reportkit_v2.py
=== Python 3.13.14, default (UTF-8 mode off)
customers  OK    ['Zoë Müller', 'Renée O’Brien', 'Åsa Lindqvist']
shop       OK    {'name': 'Café Délice', 'owner': 'José Núñez'}
report     OK    '✓ Zoë Müller (Zürich)\n'
report.txt is 28 bytes after that attempt
native     OK    'café'
ps_export  OK    'Zoë Müller'
reportkit_v2.py: 0 finding(s)

All five lines are correct. I show only the first block because the other three blocks of the matrix output are identical to it, which Step 8 turns into a test. The scanner now reports no findings. For a stricter check, make every missing encoding an exception. The flag -W error::EncodingWarning turns the warning into an error, and the driver prints it as a FAIL line:

python matrix.py -X warn_default_encoding -W error::EncodingWarning demo.py reportkit_v2

In all four runs no line prints FAIL and nothing appears on stderr, so the fixed module never relies on a default.

That leaves print(). The driver’s first line, sys.stdout.reconfigure(encoding="utf-8"), is why its output survived the old default in every block. Put the same line at the top of any script that prints non-ASCII text, or set PYTHONUTF8=1 in the environment of the job that runs it. Under UTF-8 mode, the Step 4 print check passed without either.

Step 8: Make the test suite prove it

UTF-8 mode is decided when the interpreter starts, so a test cannot flip it from the inside. The trick is to launch fresh interpreters from the test, one per mode, and compare what they print. Create test_reportkit.py:

# test_reportkit.py
import os
import subprocess
import sys
from pathlib import Path

import pytest

HERE = Path(__file__).parent
NAMES = ["Zoë Müller", "Renée O’Brien", "Åsa Lindqvist"]


@pytest.fixture(scope="session", autouse=True)
def fixtures():
    subprocess.run([sys.executable, "make_fixtures.py"], cwd=HERE, check=True, capture_output=True)


def run_demo(module, utf8, strict=False):
    env = {k: v for k, v in os.environ.items() if not k.upper().startswith("PYTHON")}
    env["PYTHONUTF8"] = utf8
    cmd = [sys.executable]
    if strict:
        cmd += ["-X", "warn_default_encoding", "-W", "error::EncodingWarning"]
    proc = subprocess.run(cmd + ["demo.py", module], cwd=HERE, env=env, capture_output=True)
    assert proc.returncode == 0, proc.stderr.decode("utf-8", "replace")
    return proc.stdout.decode("utf-8")


def run_scan(module):
    proc = subprocess.run([sys.executable, "scan_encoding.py", module], cwd=HERE, capture_output=True, text=True, encoding="utf-8")
    return proc.returncode


def test_fixed_module_behaves_the_same_in_both_modes():
    assert run_demo("reportkit_v2", "0", strict=True) == run_demo("reportkit_v2", "1", strict=True)


def test_fixed_module_reads_every_source_correctly():
    out = run_demo("reportkit_v2", "0", strict=True)
    assert "FAIL" not in out
    assert repr(NAMES) in out
    assert "'Café Délice'" in out and "'José Núñez'" in out
    assert "'✓ Zoë Müller (Zürich)\\n'" in out
    assert "'café'" in out
    assert "ps_export  OK    'Zoë Müller'" in out


def test_original_module_changes_behaviour_between_modes():
    assert run_demo("reportkit_v1", "0") != run_demo("reportkit_v1", "1")


def test_strict_warnings_flag_the_original_module():
    out = run_demo("reportkit_v1", "1", strict=True)
    assert out.count("EncodingWarning") >= 4


def test_scanner_flags_the_original_but_not_the_fixed_module():
    assert run_scan("reportkit_v1.py") == 1
    assert run_scan("reportkit_v2.py") == 0

Each test answers one question:

  • test_fixed_module_behaves_the_same_in_both_modes runs the driver with PYTHONUTF8 set to 0 and to 1, with the strict warning flags, and requires identical output. This is the guarantee you want.
  • test_fixed_module_reads_every_source_correctly checks the actual values, so “identical but identically wrong” cannot pass.
  • test_original_module_changes_behaviour_between_modes runs the original module in both modes and requires different output. It looks odd to test the broken code, but it proves the harness can see mode-dependent behavior at all. If it ever fails, your test setup has gone blind.
  • test_strict_warnings_flag_the_original_module checks that the strict flags catch at least four warnings in the original.
  • test_scanner_flags_the_original_but_not_the_fixed_module checks the scanner’s exit status, 1 for the original and 0 for the fixed module.

Run the suite under all four interpreter settings:

python matrix.py -m pytest -q test_reportkit.py
=== Python 3.13.14, default (UTF-8 mode off)
.....                                                                    [100%]
5 passed in 0.94s

=== Python 3.13.14, PYTHONUTF8=1 (UTF-8 mode on)
.....                                                                    [100%]
5 passed in 0.92s

=== Python 3.15.0rc3, default (UTF-8 mode on)
.....                                                                    [100%]
5 passed in 0.83s

=== Python 3.15.0rc3, PYTHONUTF8=0 (UTF-8 mode off)
.....                                                                    [100%]
5 passed in 0.81s

Five tests pass in each of the four runs. The outer interpreter’s mode does not matter, because every test sets the mode of its own child process. For a CI gate that covers your whole suite, run it with the strict flags:

python -X warn_default_encoding -W error::EncodingWarning -m pytest -q test_reportkit.py
.....                                                                    [100%]
5 passed in 0.94s

That passes here, which means pytest and the standard-library code these tests call raised no EncodingWarning. A larger project may have dependencies that do; the warning points at the file that omitted the encoding, so you can see whether the cause is your code or a library.

Common mistakes and gotchas

  • A preset environment hides the bugs. PYTHONUTF8=1 in your shell, your CI image or an IDE run configuration makes 3.13 behave like 3.15, and PYTHONIOENCODING=utf-8 changes the standard streams. I had both set without knowing it. Check with Get-ChildItem Env:PYTHON* before you trust any comparison.
  • Silencing the error is not fixing it. Adding errors="replace" or errors="ignore" makes a decode failure disappear by storing � or dropping the character, exactly the loss you saw in Step 1. Name the right encoding instead.
  • The flag does not reach child Python processes. Use PYTHONUTF8 when tests or tools start other interpreters (Step 5).
  • A subprocess decode error can surface as None. On Windows, check stderr when result.stdout is None (Step 4).
  • A file that opens fine is not proof of the right encoding. The old default read the UTF-8 customers file without an error and produced wrong names. Compare against known values, as the second test does.
  • Each scanner has blind spots. Run the runtime check, a static scan and Ruff together (Step 6).
  • Files that start with a BOM need a different codec. A UTF-8 file saved with a BOM, which some Windows tools add, needs encoding="utf-8-sig"; my CSV parsing tutorial shows that failure and fix.

Check the whole thing end to end

From a clean encoding-lab folder that holds the files from this tutorial, these commands should reproduce the results above:

python make_fixtures.py
python matrix.py demo.py reportkit_v1
python matrix.py demo.py reportkit_v2
python scan_encoding.py reportkit_v1.py
python scan_encoding.py reportkit_v2.py
python matrix.py -m pytest -q test_reportkit.py

You should see four result blocks for the original module that fall into two different pairs, four identical blocks for the fixed one, five scanner findings with exit status 1 and then none with exit status 0, and five passing tests in each of four runs. If the first command prints different byte values, your regional settings differ from mine; the pattern should still hold.

Where to go next

Run scan_encoding.py on your own files, then add Ruff’s PLW1514 rule to CI so new code cannot reintroduce the problem. Set PYTHONWARNDEFAULTENCODING in a CI job that runs your tests, and decide each site with the table from Step 7. Upgrade when you are ready: with every encoding named, the Python version stops mattering.

Three related tutorials on this site pair well with this one. The Python 3.15 lazy imports tutorial and the Tachyon sampling profiler tutorial cover the other two 3.15 features I have tested the same way. Once your bytes decode correctly, the invisible Unicode tutorial shows how to check which characters the text actually contains.

Tags:

Character EncodingpytestPythonPython 3.15UnicodeWindows

Share

A row of green and grey fibre broadband street cabinets on a pavement beside a fence in Iver, England
Previous Post

BT’s TalkTalk Rescue Turns Telecom Continuity Into a New Merger-Control Ground

Denmark’s oldest church book, from Holmens parish, open on a stack of books; its handwritten pages record births between 1617 and 1639
Next Post

Denmark Says 8.8 Million Population Register Records Were Pulled Through One Company’s Lawful Access

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
05 Oct
How to Use frozendict in Python 3.15 to Freeze Config and Cache Dictionary Arguments
05 Oct
Kubernetes Node Swap Turns Idle Agent Memory Into a Density Bet With No Wake-Up Test
Trending
October 5, 2026
How to Use frozendict in Python 3.15 to Freeze Config and Cache Dictionary Arguments
October 5, 2026
Kubernetes Node Swap Turns Idle Agent Memory Into a Density Bet With No Wake-Up Test
October 5, 2026
Denmark Says 8.8 Million Population Register Records Were Pulled Through One Company’s Lawful Access
October 5, 2026
How to Prepare Your Python Code for the Python 3.15 UTF-8 Default and Fix Windows Encoding Bugs
October 5, 2026
BT’s TalkTalk Rescue Turns Telecom Continuity Into a New Merger-Control Ground
October 5, 2026
Google Stops Accepting Product Bug Reports for Its Open-Source Bounty, Citing Automated Submissions

Related Posts

A laptop wrapped in a chain and padlock, illustrating least-privilege controls for AI agents.
Learning Hub

How to Secure Tool-Using AI Agents Before They Touch Production

June 8, 2026
Colorful sticky notes arranged on an office wall, symbolizing governance checklists and planning.
Learning Hub

AI Governance for Agentic Apps: A Practical Checklist for Builders

June 8, 2026
A technician connects green fiber optic cables at a data center, representing a private production inference endpoint.
Learning Hub

How to Deploy a Fine-Tuned LLM Behind a Private Production Inference Endpoint

June 8, 2026
Narrow aisle behind black supercomputer racks in a data center
Learning Hub

Kubernetes SELinux Volume Labeling: What Cluster Operators Should Audit Before v1.37

June 8, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026