How to Safely Refactor Legacy Python Code Using Characterization Tests
Learn how to safely refactor untested legacy Python code by writing characterization tests that lock in real, observed behavior before you touch a single line.
Legacy code has a reputation problem. Most developers do not fear old code because it looks ugly; they fear it because nobody can say with confidence what it actually does, or what will break if it changes. That fear causes teams to freeze: bugs and all, the code stays exactly as it is because touching it feels too risky.
Table Of Content
- What Is a Characterization Test?
- Prerequisites
- Step 1: Set Up Your Project
- Step 2: Meet the Legacy Function
- Step 3: Discover the Real Behavior, Don’t Guess It
- Step 4: Turn Your Observations Into a Characterization Test Suite
- Step 5: Refactor With the Safety Net in Place
- Step 6: Prove the Safety Net Actually Works
- Step 7: Verify an AI-Suggested Refactor Before You Trust It
- Common Mistakes and Gotchas
- How to Verify Everything Works End to End
A characterization test breaks that freeze. Instead of writing a test that says what code should do, you write a test that records what the code actually does right now, quirks included. Once that safety net is in place, you can refactor with real evidence that you did not change behavior, rather than a hope that you did not.
In this tutorial you will build a small, deliberately tangled Python function with no tests, pin down its real behavior using characterization tests, refactor it into clean, independently testable pieces, and then prove the safety net works by deliberately breaking the code and watching the tests catch it. Along the way you will also ask a local AI model to propose its own refactor and learn why you should verify that suggestion instead of trusting it outright.
What Is a Characterization Test?
The term comes from Michael Feathers’ book Working Effectively with Legacy Code. According to the definition on Wikipedia’s characterization test entry, a characterization test is “a means to describe (characterize) the actual behavior of an existing piece of software, and therefore protect existing behavior of legacy code against unintended changes via automated testing.”
That definition hides an important reversal of how most people learn to write tests. Ordinary unit tests start from a specification: you decide what the correct output should be, then assert that the code produces it. A characterization test starts from observation: you run the code, see what it actually outputs, and assert that value. As the same source puts it, “when creating a characterization test, one must observe what outputs occur for a given set of inputs.”
That is not a shortcut or a lesser form of testing. It is the right tool for a specific job: making it safe to change code you did not write and do not fully trust, without accidentally taking on the separate (and much bigger) job of deciding what the code should have done all along. If a characterization test later turns up a genuine bug, you fix the bug on purpose, as a visible, reviewed decision, not as a side effect of a refactor.
Prerequisites
- Python 3.10 or later. This tutorial was built and tested on Python 3.13.14.
- pip, to install pytest.
- Basic familiarity with Python functions. No prior testing experience is required; this tutorial explains pytest as it goes.
- Optional: Ollama installed locally with a model pulled (this tutorial uses
qwen3.5:4b), only needed for the final, optional AI-assisted step.
Step 1: Set Up Your Project
Create a fresh project folder and install pytest, the test runner you will use throughout this tutorial.
mkdir chartest-demo
cd chartest-demo
python -m venv venv
venv\Scripts\activate # Windows
# source venv/bin/activate # macOS/Linux
pip install pytest
Confirm the install worked:
python -m pytest --version
Expected output (versions may differ slightly):
pytest 9.1.1
Step 2: Meet the Legacy Function
Create a file named legacy_pricing.py with the function below. Imagine you inherited this from a codebase with no tests and no one left on the team who wrote it, a common definition of “legacy code” that has nothing to do with the age of the code and everything to do with the absence of a safety net.
audit_log = []
notifications = []
def process_order(customer_type, subtotal, item_count, is_first_order):
if subtotal < 0 or item_count < 0:
raise ValueError("subtotal and item_count must not be negative")
if customer_type == "vip":
rate = 0.20
elif customer_type == "member":
rate = 0.10
else:
rate = 0.0
if item_count >= 10:
rate += 0.05
if is_first_order:
rate += 0.03
if rate > 0.25:
rate = 0.25
discount = subtotal * rate
total = subtotal - discount
total = int(total * 100) / 100
audit_log.append(f"order total={total} rate={rate}")
if total > 100:
notifications.append(f"large order notice: {total}")
return total
Read through it once. A few things should stand out as risky to change blindly:
- Discount rate logic, rounding, logging, and a notification side effect are all tangled together in one function.
- The final line,
int(total * 100) / 100, truncates to two decimal places rather than rounding. That is easy to misread as a rounding bug on a quick skim. - There is a rate cap at 0.25 whose interaction with the other bonuses is not obvious just from reading the code.
None of that is necessarily wrong. It might be exactly the behavior the business relies on. The point of a characterization test is that you do not have to answer that question before you can safely refactor: you just have to preserve whatever the answer currently is.
Before you change anything, make a copy of this file as legacy_pricing_original.py. You will not need it until Step 7, but it is easy to forget once you start editing legacy_pricing.py directly.
Step 3: Discover the Real Behavior, Don’t Guess It
Here is the core characterization testing technique, and it feels backwards the first time you do it: instead of writing what you think the output should be, you write an assertion you know is wrong on purpose, run it, and let pytest’s failure message tell you the real value.
Create test_discover.py:
from legacy_pricing import process_order
def test_guest_small_order():
result = process_order("guest", 50.0, 2, False)
assert result == 0
def test_member_first_order():
result = process_order("member", 200.0, 3, True)
assert result == 0
def test_vip_bulk_order():
result = process_order("vip", 300.0, 12, False)
assert result == 0
def test_vip_bulk_first_order_caps_rate():
result = process_order("vip", 300.0, 12, True)
assert result == 0
Every assertion claims the result is 0, which is almost certainly false for a paid order. That is intentional. Run it:
python -m pytest test_discover.py -v
Real, captured output from this exact code:
test_discover.py::test_guest_small_order FAILED [ 25%]
test_discover.py::test_member_first_order FAILED [ 50%]
test_discover.py::test_vip_bulk_order FAILED [ 75%]
test_discover.py::test_vip_bulk_first_order_caps_rate FAILED [100%]
================================== FAILURES ===================================
___________________________ test_guest_small_order ____________________________
def test_guest_small_order():
result = process_order("guest", 50.0, 2, False)
> assert result == 0
E assert 50.0 == 0
___________________________ test_member_first_order ___________________________
E assert 174.0 == 0
_____________________________ test_vip_bulk_order _____________________________
E assert 225.0 == 0
_____________________ test_vip_bulk_first_order_caps_rate _____________________
E assert 225.0 == 0
=========================== short test summary info ===========================
4 failed in 0.05s
pytest’s assertion rewriting shows you both sides of every failed comparison, so the real values (50.0, 174.0, 225.0, 225.0) are right there in the failure output. You did not calculate those by hand and you did not have to trust a docstring or a comment. You watched the code tell you what it does.
Notice something interesting: the last two tests both return 225.0, even though one of them includes the first-order bonus and the other does not. That is the 0.25 rate cap in action: a VIP customer buying 12 items already hits the cap at 20% + 5% = 25%, so the extra 3% first-order bonus has no visible effect at all in this case. That is exactly the kind of non-obvious interaction a characterization test is good at pinning down before you touch the code.
Step 4: Turn Your Observations Into a Characterization Test Suite
Now flip those wrong assertions into correct ones, using the exact values pytest just showed you. Create test_characterize.py:
from legacy_pricing import process_order
def test_guest_small_order():
result = process_order("guest", 50.0, 2, False)
assert result == 50.0
def test_member_first_order():
result = process_order("member", 200.0, 3, True)
assert result == 174.0
def test_vip_bulk_order():
result = process_order("vip", 300.0, 12, False)
assert result == 225.0
def test_vip_bulk_first_order_caps_rate():
result = process_order("vip", 300.0, 12, True)
assert result == 225.0
def test_rejects_negative_subtotal():
import pytest
with pytest.raises(ValueError):
process_order("guest", -10.0, 1, False)
def test_member_order_truncates_not_rounds():
result = process_order("member", 23.32, 1, False)
assert result == 20.98
That last test deserves a closer look, because it is the one that will matter most later. With a subtotal of 23.32 and a 10% member discount, the exact total is 20.988. Truncating to two decimals gives 20.98; rounding would give 20.99. Picking an input where those two behaviors disagree is what makes this test worth having: it does not just check that the function returns a number, it locks in which rounding strategy the legacy code actually uses.
You can delete test_discover.py now; it was scaffolding for finding the numbers, not a test worth keeping. Run the real suite:
python -m pytest test_characterize.py -v
Real captured output:
test_characterize.py::test_guest_small_order PASSED [ 16%]
test_characterize.py::test_member_first_order PASSED [ 33%]
test_characterize.py::test_vip_bulk_order PASSED [ 50%]
test_characterize.py::test_vip_bulk_first_order_caps_rate PASSED [ 66%]
test_characterize.py::test_rejects_negative_subtotal PASSED [ 83%]
test_characterize.py::test_member_order_truncates_not_rounds PASSED [100%]
6 passed in 0.02s
This green baseline is your safety net. Commit it before you change a single line of the function it protects. If you refactor first and write tests after, you cannot tell whether a later red test means you broke something or means the test itself is new and still being written.
Step 5: Refactor With the Safety Net in Place
With the tests in place, split the tangled function into small, single-purpose pieces. Replace the contents of legacy_pricing.py with this:
audit_log = []
notifications = []
RATE_BY_CUSTOMER_TYPE = {"vip": 0.20, "member": 0.10}
BULK_ITEM_THRESHOLD = 10
BULK_BONUS = 0.05
FIRST_ORDER_BONUS = 0.03
MAX_DISCOUNT_RATE = 0.25
def calculate_discount_rate(customer_type, item_count, is_first_order):
rate = RATE_BY_CUSTOMER_TYPE.get(customer_type, 0.0)
if item_count >= BULK_ITEM_THRESHOLD:
rate += BULK_BONUS
if is_first_order:
rate += FIRST_ORDER_BONUS
return min(rate, MAX_DISCOUNT_RATE)
def apply_discount(subtotal, rate):
discount = subtotal * rate
total = subtotal - discount
# Truncates rather than rounds, e.g. 20.988 becomes 20.98, not 20.99.
# This is the legacy function's original behavior; the characterization
# tests exist specifically to protect it during this refactor.
return int(total * 100) / 100
def process_order(customer_type, subtotal, item_count, is_first_order):
if subtotal < 0 or item_count < 0:
raise ValueError("subtotal and item_count must not be negative")
rate = calculate_discount_rate(customer_type, item_count, is_first_order)
total = apply_discount(subtotal, rate)
audit_log.append(f"order total={total} rate={rate}")
if total > 100:
notifications.append(f"large order notice: {total}")
return total
The discount-rate logic and the rounding logic are now separate, independently readable functions with names that say what they do. process_order is left as a thin orchestrator: validate, calculate, apply, log, notify, return.
Crucially, test_characterize.py has not changed at all. Run it again exactly as it is:
python -m pytest test_characterize.py -v
Real captured output:
test_characterize.py::test_guest_small_order PASSED [ 16%]
test_characterize.py::test_member_first_order PASSED [ 33%]
test_characterize.py::test_vip_bulk_order PASSED [ 50%]
test_characterize.py::test_vip_bulk_first_order_caps_rate PASSED [ 66%]
test_characterize.py::test_rejects_negative_subtotal PASSED [ 83%]
test_characterize.py::test_member_order_truncates_not_rounds PASSED [100%]
6 passed in 0.02s
Same six tests, same six passes, completely different internal structure. That is the entire point: the tests only look at what process_order returns (and, for the negative-input case, what it raises), never at how it gets there. That is what lets you rearrange the inside freely.
Step 6: Prove the Safety Net Actually Works
A test suite that only ever passes is not proof of anything by itself; you also need to see it fail when it should. Deliberately introduce a plausible-looking “cleanup.” In apply_discount, swap the truncation for what looks like a bug fix:
def apply_discount(subtotal, rate):
discount = subtotal * rate
total = subtotal - discount
# "Fixed" to round properly instead of truncating.
return round(total, 2)
This change is easy to justify to yourself in the moment: rounding looks more correct than truncating, and most of your test cases will not even notice the difference. Run the suite anyway:
python -m pytest test_characterize.py -v
Real captured output:
test_characterize.py::test_guest_small_order PASSED [ 16%]
test_characterize.py::test_member_first_order PASSED [ 33%]
test_characterize.py::test_vip_bulk_order PASSED [ 50%]
test_characterize.py::test_vip_bulk_first_order_caps_rate PASSED [ 66%]
test_characterize.py::test_rejects_negative_subtotal PASSED [ 83%]
test_characterize.py::test_member_order_truncates_not_rounds FAILED [100%]
================================== FAILURES ===================================
___________________ test_member_order_truncates_not_rounds ____________________
def test_member_order_truncates_not_rounds():
result = process_order("member", 23.32, 1, False)
> assert result == 20.98
E assert 20.99 == 20.98
1 failed, 5 passed in 0.04s
Five of six tests do not notice anything changed, because for most inputs truncating and rounding produce the same number. Only the test you deliberately built around a fractional edge case catches it, with a precise, honest diff: the code now returns 20.99 where the characterization suite says it must return 20.98.
This is exactly why “just add a few obvious test cases” is not the same guarantee as characterization testing: it is easy to write tests that all happen to land on inputs where two subtly different implementations agree. Revert the change before continuing:
def apply_discount(subtotal, rate):
discount = subtotal * rate
total = subtotal - discount
return int(total * 100) / 100
Run the suite once more to confirm you are back to a clean baseline (6 passed) before moving on.
Step 7: Verify an AI-Suggested Refactor Before You Trust It
AI coding tools are genuinely useful for exactly this kind of mechanical refactor, but “the code compiles and looks reasonable” is not the same as “the behavior is unchanged.” Your characterization suite is what closes that gap. This step is optional and uses a local model through Ollama’s HTTP API, so no data leaves your machine and no API key is required.
With Ollama running and a model pulled (ollama pull qwen3.5:4b), save this as ask_ollama_refactor.py, pointing it at your original, untouched legacy function:
import json
import urllib.request
with open("legacy_pricing_original.py", "r", encoding="utf-8") as f:
original_source = f.read()
prompt = f"""You are refactoring a legacy Python function. Split process_order into
small, pure helper functions (e.g. one for the discount rate, one for applying it),
while preserving its exact external behavior, including any rounding/truncation
quirks. Return ONLY the refactored Python source code, no explanation.
{original_source}"""
payload = {"model": "qwen3.5:4b", "prompt": prompt, "stream": False, "think": False}
req = urllib.request.Request(
"http://localhost:11434/api/generate",
data=json.dumps(payload).encode("utf-8"),
headers={"Content-Type": "application/json"},
)
with urllib.request.urlopen(req, timeout=120) as resp:
result = json.loads(resp.read().decode("utf-8"))
print(result["response"])
stream: false asks Ollama’s /api/generate endpoint for one complete response instead of a series of streamed chunks, and think: false skips the model’s internal reasoning trace since you only want the final code. Run it:
python ask_ollama_refactor.py
This is the real response captured from a local qwen3.5:4b run while writing this tutorial. Model output is not deterministic, so if you run this yourself you will likely get different code, possibly correct code. That unpredictability is itself part of the lesson: you cannot assume any single response is safe just because a previous one was.
audit_log = []
notifications = []
def get_discount_rate(customer_type, item_count, is_first_order) -> float:
rate = 0.0
if customer_type == "vip":
rate += 0.20
elif customer_type == "member":
rate += 0.10
if item_count >= 10:
rate += 0.05
if is_first_order:
rate += 0.03
return max(rate, 0.0)
def calculate_final_amount(subtotal, discount_rate): -> int:
global audit_log
global notifications
total = subtotal - (subtotal * discount_rate)
# Round to nearest cent using truncation logic equivalent
total_int_cents = int(total * 100)
total_rounded = round(total_int_cents / 100.0, 2)
audit_log.append(f"order total={total_rounded} rate={discount_rate}")
if total_rounded > 100:
notifications.append(f"large order notice: {total_rounded}")
return int(total * 100 / 100)
Save the printed response into a new file, ai_suggestion.py, so you can check it without touching your working legacy_pricing.py yet. Do not paste it in and assume it works because it looks reasonable. Try to actually use it first:
python ai_suggestion.py
Real captured output:
File "ai_suggestion.py", line 22
def calculate_final_amount(subtotal, discount_rate): -> int:
^^
SyntaxError: invalid syntax
The return-type annotation is in the wrong place; it belongs before the colon, as def calculate_final_amount(subtotal, discount_rate) -> int:. That much is an easy catch: the file will not even import. But a syntax error is the friendly kind of AI mistake, the kind that stops you before you ship it. Fix just that one line. Then, before overwriting anything, make a copy of your working legacy_pricing.py from Step 5 (call it legacy_pricing_working.py) so you have something to restore from, copy the syntax-corrected suggestion over legacy_pricing.py, and run the characterization suite before assuming the rest is fine:
python -m pytest test_characterize.py -v
Real captured output, against the syntax-corrected version of the exact code above:
test_characterize.py::test_guest_small_order PASSED [ 16%]
test_characterize.py::test_member_first_order PASSED [ 33%]
test_characterize.py::test_vip_bulk_order PASSED [ 50%]
test_characterize.py::test_vip_bulk_first_order_caps_rate FAILED [ 66%]
test_characterize.py::test_rejects_negative_subtotal PASSED [ 83%]
test_characterize.py::test_member_order_truncates_not_rounds FAILED [100%]
================================== FAILURES ===================================
_____________________ test_vip_bulk_first_order_caps_rate _____________________
def test_vip_bulk_first_order_caps_rate():
result = process_order("vip", 300.0, 12, True)
> assert result == 225.0
E assert 216 == 225.0
___________________ test_member_order_truncates_not_rounds ____________________
def test_member_order_truncates_not_rounds():
result = process_order("member", 23.32, 1, False)
> assert result == 20.98
E assert 20 == 20.98
2 failed, 4 passed in 0.04s
Two real, distinct regressions, caught before either could reach production:
- The 0.25 rate cap silently disappeared.
get_discount_ratereturnsmax(rate, 0.0), a floor with no ceiling. Nothing in the code looks obviously wrong on a skim, and a VIP customer with a large, first-time order would quietly get a bigger discount than the business ever agreed to. - The rounding got worse, not better. Look closely at the final line:
int(total * 100 / 100). Multiplication and division have the same precedence and evaluate left to right, sototal * 100 / 100is mathematically justtotalagain, and wrapping that inint()truncates to a whole number. The function returns whole-dollar totals, discarding every cent, not the two-decimal-place truncation the audit-log line right above it is busy calculating correctly and using for the log and notification, but not for the number actually returned to the caller.
Neither bug would show up by reading the code once and nodding along; both look like reasonable refactors at a glance. They only show up because six characterization tests, built from real observed behavior instead of assumptions, disagreed with the new code and said so precisely: 216 instead of 225.0, 20 instead of 20.98. That precision is what makes the failure actionable instead of just alarming: you know exactly what changed and can decide, deliberately, whether to fix the suggestion or discard it. The lesson is not “AI refactors are unreliable.” It is that the same verification step applies whether a human or a model wrote the change, AI-generated code included, and you already built the tool that makes that verification cheap and fast.
Common Mistakes and Gotchas
- Designing expected values instead of observing them. If you calculate what you think the answer should be and write that into the assertion, you are testing your assumptions about the code, not the code itself. Always let the deliberately-wrong-assertion trick (Step 3) show you the real value first.
- Testing implementation details instead of outputs. These tests only look at what
process_orderreturns or raises. If you had instead asserted on internal variable names or the exact structure ofaudit_logentries, the refactor in Step 5 could have broken tests for no behavioral reason at all, a classic source of brittle tests. - Only characterizing the happy path. Bugs cluster at boundaries. The rate-cap interaction in Step 3 and the truncation-versus-rounding case in Step 4 were both found by deliberately picking edge-of-behavior inputs, not typical ones.
- Refactoring before committing the characterization tests. If your test suite and your refactor land in the same uncommitted change, a red test cannot tell you whether you introduced a real regression or just have not finished editing the tests yet. Get a clean, committed green baseline first.
- Treating characterization tests as permanent. They are scaffolding for a specific, risky moment: the transition from untested to tested code. Once you understand a piece of code well enough to say what it should do, consider writing intention-revealing tests alongside or in place of the characterization ones, especially for any behavior you now know is an actual bug rather than a real requirement.
How to Verify Everything Works End to End
To confirm your own setup is solid, run through this checklist in order:
- Run
python -m pytest test_characterize.py -vagainst the refactoredlegacy_pricing.pyfrom Step 5 and confirm all 6 tests pass. - Temporarily reintroduce the
round(total, 2)change from Step 6 and confirm exactly one test fails, with the messageassert 20.99 == 20.98. If a different test fails, or none do, your test inputs are not actually exercising the truncation behavior. - Revert the change and confirm you are back to
6 passed. - If you tried the optional AI step, run the characterization suite against whatever code your model actually suggested before you keep any of it. A pass is not automatically safe to trust either; a fail, like the two shown in Step 7, is the safety net doing its job. Either way, read the result and decide deliberately rather than assuming the suggestion is fine because it compiled.
From here, a natural next step is property-based testing, which generates many inputs automatically instead of the handful you picked by hand, useful once you want more confidence that your characterization tests cover the edge cases that matter. You can also apply this same discover-then-lock-in technique to a real function in a codebase you maintain: pick one small, untested, slightly scary function, characterize it this week, and refactor it once the tests are green.








No Comment! Be the first one.