TRENDING
Rows of identical brass-colored apartment mailboxes with small locks and name labels along an orange corridor wall
October 9, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
Street-level upward view of the Monetary Authority of Singapore building and neighbouring office towers under a pale sky
October 9, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
Cast-iron late Qing dynasty coin minting press with a large flywheel, displayed in a museum case
October 9, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google
Rows of closed oak library card catalog drawers, each with a brass pull and a blank label holder
October 9, 2026
How to Encrypt PII in Python and Keep It Searchable With Blind Indexes
Close-up of a vintage Western Electric manual telephone switchboard with orange lamps, red patch cords plugged into jacks, a rotary dial and a black handset
October 9, 2026
Microsoft’s Agent Lightning v1.0 Turns Agent Training Into a Sample-Accounting Problem
09 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Two orange safety relief valves on grey pressure vessels in an industrial plant
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
Yellow diamond-shaped merging traffic warning sign showing a side road joining a main road
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A lugworm lying on wet sand and mud at low tide
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 232 Posts
News 234 Posts
Learning Hub 204 Posts
Home/Learning Hub/How to Debug Python Code With a Local AI Agent Using Ollama and pytest
Learning Hub

How to Debug Python Code With a Local AI Agent Using Ollama and pytest

Build a local Ollama debugging agent, watch it misdiagnose a real money bug with too little context, then fix it once you give it a failing test, real tools, and a forced verification loop.

August 17, 2026 15 Min Read
44

A teammate reports that a $10 dinner bill, split three ways in your expense-tracking app, doesn’t add up: the app charged three people $3.33 each, which totals $9.99, a penny short of the actual bill. Your existing test suite is green. Nothing looks obviously wrong in the code. This is exactly the kind of bug where handing the problem to an AI coding agent sounds appealing, and exactly the kind of bug where a careless agent can make things worse instead of better.

Table Of Content

  • What You Will Build
  • Prerequisites
  • How AI-Assisted Debugging Actually Works
  • Step 1: Set Up a Project With a Real Bug
  • The Expense Splitter
  • The Existing Test Suite (Green, But Incomplete)
  • Step 2: Reproduce the Bug With a Failing Test
  • Confirming the Bug by Hand
  • Writing a Regression Test
  • Running pytest to Capture the Failure
  • Step 3: Provide Debugging Context to the Agent
  • Attempt 1: The Bare Error Message Is Not Context
  • Building a Minimal Debugging Agent With Ollama Tool Calling
  • Attempt 2: Giving the Agent Real Tools (and Watching It Skip Verification)
  • Step 4: Apply and Verify the Fix
  • Why the First Agent Run Made Things Worse
  • Forcing a Verification Loop
  • Confirming the Fix Independently
  • Common Mistakes to Avoid
  • Trusting the Agent’s Summary Instead of the Test Output
  • Letting the Agent Weaken a Test to Make It Pass
  • Not Forcing Verification Between Edits
  • Giving a Bare Error Message Instead of a Reproducible Failure
  • How to Confirm It All Works End to End
  • Next Steps

This tutorial builds a real, working “debug this bug” workflow using a local AI agent, then runs it twice on a genuine bug to show what actually happens. The first run gives the agent almost no context and watches it recommend a fix that would have hidden the bug rather than corrected it. The second run gives the agent real tools, a failing test, and an explicit instruction to verify its own work, and still catches the agent thrashing through broken rewrites before a stricter loop forces it to check itself. Every command, every piece of model output, and every test result below is real: captured from an actual local Ollama model working against an actual buggy Python file, not paraphrased or invented for the article.

Before starting, two terms matter. An AI coding agent is different from a chatbot: instead of only replying with text, it can call tools, functions you define and hand to the model, like “read this file” or “run the tests,” and use their real output to decide what to do next. That request-a-tool, read-the-result, decide-again loop is called tool calling, and it’s what turns a model that can only talk about your code into one that can actually look at it.

What You Will Build

  • A small Python “expense splitter” module with a real, reproducible money bug
  • A regression test that pins the bug down before you try to fix it
  • A first debugging attempt using only a bare error message, and a look at why the agent’s proposed fix would have been a mistake
  • A minimal local debugging agent built from scratch with Ollama’s tool-calling API, giving the model real read_file, write_file, and run_pytest tools
  • A second debugging attempt showing the agent diagnose the bug correctly, but also showing what happens when nothing forces it to check its own work
  • A corrected agent loop that requires verification after every edit, run for real until the entire test suite passes

Prerequisites

  • Windows, macOS, or Linux. This tutorial was written and tested on Windows with Python 3.13.14; nothing here depends on the OS beyond the virtual environment activation command.
  • Python 3.10 or newer, and enough command-line comfort to create a virtual environment and run a script.
  • Basic familiarity with pytest: you write functions starting with test_ containing assert statements, and pytest runs them and reports which passed and which failed. If you’ve never used it, that one sentence is enough to follow along.
  • Comfortable reading a Python traceback well enough to tell which line raised an error.
  • Ollama installed and running locally, with at least one tool-calling-capable model pulled. This tutorial uses qwen3.5:4b (about 3.4 GB): ollama pull qwen3.5:4b. Ollama version used here is 0.32.9; check yours with ollama --version.
  • No API keys and no cloud account. Everything below talks to Ollama on localhost.

How AI-Assisted Debugging Actually Works

“AI-assisted debugging” just means pairing with an AI agent instead of hunting a bug down entirely by hand with tools like pdb or scattered print() calls. Done carelessly, it’s just as likely to produce a confident, wrong answer as a correct one, which is the whole reason this tutorial runs the same bug through the process twice. The workflow that actually works has four parts, and each one exists to remove a specific way this can go wrong:

  1. Set up a workspace the agent can safely read and write in. An agent needs real tools, not just a chat window, to gather its own evidence instead of guessing from a paraphrase.
  2. Reproduce the bug with a failing test. “The totals are wrong sometimes” is not something an agent (or a human) can verify against. A failing test is a precise, checkable target: the bug is fixed exactly when that test, and every other test, passes.
  3. Give the agent real context, not just the error. A bare error message describes a symptom. The failing test, the traceback, and the relevant source code describe the actual problem. You’ll see below exactly how much that difference matters.
  4. Apply the fix, then verify it, don’t just apply it. An agent that edits a file and stops has not fixed anything until something re-runs the tests and confirms it. This is the step most naive agent loops skip, and it’s where the more serious failure in this tutorial happens.

Step 1: Set Up a Project With a Real Bug

The Expense Splitter

Create a project folder, a virtual environment, and install pytest and requests:

mkdir ai-debug-tutorial
cd ai-debug-tutorial
python -m venv venv
venv\Scripts\activate   # on Linux/macOS: source venv/bin/activate
pip install pytest requests

This was written against pytest 9.1.1 and requests on Python 3.13.14. Now save the module you’ll be debugging as splitter.py:

"""Split a shared expense evenly among a group of people."""


def split_expense(total, people):
    """Split `total` dollars evenly among `people` names.

    Returns a dict mapping each person's name to the amount they owe.
    """
    if not people:
        raise ValueError("Need at least one person to split an expense")

    share = round(total / len(people), 2)
    return {person: share for person in people}

This is the kind of function that looks obviously correct on a read-through: divide the total by the number of people, round to the nearest cent, done. That’s exactly what makes the bug worth teaching with.

The Existing Test Suite (Green, But Incomplete)

Save this as test_splitter.py:

from splitter import split_expense


def test_split_expense_three_people():
    result = split_expense(30.00, ["Amir", "Sam", "Lee"])
    assert result == {"Amir": 10.00, "Sam": 10.00, "Lee": 10.00}


def test_split_expense_two_people():
    result = split_expense(50.00, ["Amir", "Sam"])
    assert result == {"Amir": 25.00, "Sam": 25.00}


def test_split_expense_requires_people():
    try:
        split_expense(20.00, [])
    except ValueError:
        pass
    else:
        raise AssertionError("Expected a ValueError for an empty people list")

Run it:

python -m pytest -v
test_splitter.py::test_split_expense_three_people PASSED                 [ 33%]
test_splitter.py::test_split_expense_two_people PASSED                   [ 66%]
test_splitter.py::test_split_expense_requires_people PASSED              [100%]

============================== 3 passed in 0.03s ==============================

All green. This matters: $30 split three ways and $50 split two ways both divide evenly, so these tests can’t catch the bug. This is a realistic starting point. Plenty of real bugs survive an honestly-written, fully-passing test suite because the suite happens to only exercise the inputs where the code behaves correctly.

Step 2: Reproduce the Bug With a Failing Test

Confirming the Bug by Hand

Before writing anything, check the bug report is real:

python -c "
from splitter import split_expense
result = split_expense(10.00, ['Amir', 'Sam', 'Lee'])
print('Shares:', result)
print('Sum:', sum(result.values()))
"
Shares: {'Amir': 3.33, 'Sam': 3.33, 'Lee': 3.33}
Sum: 9.99

Confirmed. $10 split three ways charges three people $3.33 each, and $3.33 times three is $9.99, a cent less than the original bill. In a real expense app, that missing cent has to come from somewhere: either the group quietly under-pays every time a bill doesn’t divide evenly, or someone’s books stop reconciling.

Writing a Regression Test

A verbal description of a bug isn’t something an agent (or you, six months from now) can check against. Turn it into a test that fails for exactly the reported reason. Add this to test_splitter.py:

def test_split_expense_total_matches_original():
    result = split_expense(10.00, ["Amir", "Sam", "Lee"])
    assert sum(result.values()) == 10.00

Running pytest to Capture the Failure

python -m pytest -v
test_splitter.py::test_split_expense_three_people PASSED                 [ 25%]
test_splitter.py::test_split_expense_two_people PASSED                   [ 50%]
test_splitter.py::test_split_expense_requires_people PASSED              [ 75%]
test_splitter.py::test_split_expense_total_matches_original FAILED       [100%]

================================== FAILURES ===================================
__________________ test_split_expense_total_matches_original __________________

    def test_split_expense_total_matches_original():
        result = split_expense(10.00, ["Amir", "Sam", "Lee"])
>       assert sum(result.values()) == 10.00
E       AssertionError: assert 9.99 == 10.0
E        +  where 9.99 = sum(dict_values([3.33, 3.33, 3.33]))

test_splitter.py:25: AssertionError
================== 1 failed, 3 passed in 0.08s (output trimmed) ===============

This is the real deliverable of Step 2: one failing test, with a full traceback, sitting alongside three passing ones. “Fixed” now has an exact, mechanical definition: all four tests pass, and none of the three that already passed stop passing.

Step 3: Provide Debugging Context to the Agent

Attempt 1: The Bare Error Message Is Not Context

The most tempting shortcut is to copy the assertion error into a chat prompt and ask for a fix, without giving the agent the source code, the test, or any way to look at either. Here’s exactly that, sent to a local qwen3.5:4b model over Ollama’s chat API:

import requests

response = requests.post(
    "http://localhost:11434/api/chat",
    json={
        "model": "qwen3.5:4b",
        "messages": [
            {"role": "system", "content": "You are a Python debugging assistant. Diagnose bugs and propose fixes."},
            {"role": "user", "content": "My pytest suite has a failing test:\n\nAssertionError: assert 9.99 == 10.0\n\nFix the bug."},
        ],
        "stream": False,
        "think": False,
    },
    timeout=120,
)
print(response.json()["message"]["content"])

The real response (trimmed for length, nothing added or altered):

The error `assertionerror: assert 9.99 == 10.0` indicates that your code is
performing floating-point arithmetic which resulted in an unexpected value,
or you are comparing two floats directly using equality (`==`). Floating-point
numbers cannot be precisely represented in binary IEEE-754 format, leading to
tiny rounding errors (like `9.989999...`) when they should theoretically
equal `10.0`.

[... several code examples omitted ...]

**Final Recommendation**: If you're seeing `9.99` instead of `10`, it's likely
due to rounding errors in binary arithmetic or an intermediate calculation
step. The quickest fix is replacing your assertion line with `.approx()`.

This is a plausible-sounding, wrong diagnosis, and it matters why. pytest’s own documentation confirms pytest.approx() is a real, useful tool built specifically for one problem: binary floating-point noise, like 0.1 + 0.2 not exactly equalling 0.3. That is not this bug. This bug is a whole missing cent, caused by rounding each person’s share independently before summing, and it is completely reproducible, not noise. Proof: pytest’s own default tolerance for approx() is a relative 1e-6, tight enough that it would still correctly catch this exact failure:

python -c "
import pytest
print(9.99 == pytest.approx(10.0))              # default tolerance
print(9.99 == pytest.approx(10.0, rel=1e-2))     # the model's actual suggestion
"
False
True

The model’s actual recommendation loosened that tolerance to rel=1e-2 (a 1% margin) specifically to make this failure disappear. Applied for real, that “fix” would not touch the buggy line in splitter.py at all. It would edit the test to stop noticing that a customer got undercharged by a cent, which is worse than not fixing the bug: the test suite would go green while quietly certifying broken money math. This is the single most important failure mode to watch for when an agent “fixes” a failing test: check whether it changed the code, or just changed the test.

Building a Minimal Debugging Agent With Ollama Tool Calling

The fix is giving the agent the same three things you’d want as a human debugging this: the ability to read the relevant files, the ability to re-run the tests, and the ability to write a fix. In Ollama’s chat API, you describe each ability as a tool, a JSON Schema naming the function and its parameters, and pass the list of tools alongside your messages. When the model wants to use one, it replies with a tool_calls field instead of (or alongside) plain text; you run the real Python function yourself and feed the result back as a message with "role": "tool". This request/response shape is documented directly in Ollama’s own API reference. Save this as agent.py:

import json
import subprocess

import requests

OLLAMA_URL = "http://localhost:11434/api/chat"
MODEL = "qwen3.5:4b"

TOOLS = [
    {
        "type": "function",
        "function": {
            "name": "read_file",
            "description": "Read the full contents of a file in the project directory.",
            "parameters": {
                "type": "object",
                "properties": {
                    "path": {"type": "string", "description": "Relative file path, e.g. splitter.py"},
                },
                "required": ["path"],
            },
        },
    },
    {
        "type": "function",
        "function": {
            "name": "run_pytest",
            "description": "Run the project's pytest suite and return the raw output.",
            "parameters": {"type": "object", "properties": {}},
        },
    },
    {
        "type": "function",
        "function": {
            "name": "write_file",
            "description": "Overwrite a file in the project directory with new contents.",
            "parameters": {
                "type": "object",
                "properties": {
                    "path": {"type": "string"},
                    "contents": {"type": "string"},
                },
                "required": ["path", "contents"],
            },
        },
    },
]


def read_file(path):
    with open(path, "r", encoding="utf-8") as fh:
        return fh.read()


def write_file(path, contents):
    with open(path, "w", encoding="utf-8") as fh:
        fh.write(contents)
    return f"wrote {len(contents)} bytes to {path}"


def run_pytest():
    result = subprocess.run(
        ["venv/Scripts/python.exe", "-m", "pytest", "-v"],
        capture_output=True, text=True, timeout=60,
    )
    return (result.stdout + result.stderr)[-4000:]


DISPATCH = {"read_file": read_file, "run_pytest": run_pytest, "write_file": write_file}

SYSTEM_PROMPT = """You are a Python debugging agent working inside a real project directory.
You have tools to read files, run the test suite, and write fixed files.
Always read the failing test and the source file it exercises before proposing a fix.
Do not weaken or delete a test to make it pass. Fix the underlying code so the
existing behavior stays correct and the new test passes for the right reason.
When you are confident in a fix, use write_file to apply it, then use run_pytest
to confirm the entire suite passes."""

USER_PROMPT = """Running pytest in this project produces one failure:

test_splitter.py::test_split_expense_total_matches_original FAILED

Find the failing test, read the relevant source, diagnose the root cause, fix it,
and verify the full suite passes."""


def main():
    messages = [
        {"role": "system", "content": SYSTEM_PROMPT},
        {"role": "user", "content": USER_PROMPT},
    ]
    for turn in range(8):
        resp = requests.post(OLLAMA_URL, json={
            "model": MODEL, "messages": messages, "tools": TOOLS,
            "stream": False, "think": False,
        }, timeout=180)
        message = resp.json()["message"]
        messages.append(message)
        tool_calls = message.get("tool_calls")
        if not tool_calls:
            break
        for call in tool_calls:
            name = call["function"]["name"]
            args = call["function"]["arguments"]
            if isinstance(args, str):
                args = json.loads(args)
            result = DISPATCH[name](**args)
            messages.append({"role": "tool", "content": str(result)})


if __name__ == "__main__":
    main()

Attempt 2: Giving the Agent Real Tools (and Watching It Skip Verification)

Running this against the same buggy splitter.py produces a genuinely better start. The agent reads test_splitter.py, then splitter.py, then correctly diagnoses the root cause in its own words:

Now I understand the issue. [...] there's a loss that needs to be redistributed
to ensure the total always matches the original amount.

The fix should adjust at least one person's share by the difference between the
target sum and actual sum. A common approach is to give extra cents to someone
(typically the last or first person) until the totals add up correctly.

That diagnosis is correct. What happens next is the actual lesson of this section. Nothing in the prompt required the agent to check its work before moving on, so it didn’t. It called write_file seven times in a row, each attempt visibly worse than the last, without a single run_pytest call in between:

Turn 4: writes a version calling `result.values()[i]`, which fails immediately
        because dict_values objects don't support indexing.
Turn 5: "That was wrong. Let me rewrite this cleanly" [rewrites, still broken]
Turn 6: "Wait, that approach doesn't preserve the order correctly" [rewrites again]
Turn 7: introduces `decimals=2` as a keyword argument that doesn't exist
Turn 8: runs out of the 8-turn budget mid-function, no return statement

The final state of splitter.py after that run was not merely still buggy. It was syntactically incomplete, a function with no return statement at all, actively worse than the file it started from. This is a genuine, common failure mode for agent loops, not a contrived example: an agent that can edit files but is never forced to check its edits will keep editing, because nothing tells it to stop or that it’s wrong.

Step 4: Apply and Verify the Fix

Why the First Agent Run Made Things Worse

“Apply the fix” and “the fix is applied” are not the same claim. The first run applied seven fixes and verified zero of them. The corrected version of the system prompt adds exactly one new rule, aimed directly at that gap:

You MUST call run_pytest immediately after every write_file call, before doing
anything else. Never call write_file twice in a row without a run_pytest in
between. Look at the run_pytest result: if any test still fails, diagnose that
specific failure and try again. Only stop once run_pytest shows every test
passing.

Forcing a Verification Loop

Restoring splitter.py to its original buggy state and re-running the agent with that updated prompt (and a slightly larger 12-turn budget, since verifying takes extra turns) produces a real, working session. It reads the test and source first, diagnoses the same root cause, then writes a first attempt that is itself still buggy:

result[people[i]] += round(remainder / 100.0, 2) if not result.values()[i] == base_share else None
    # Simplify approach - use a more reliable method

This time, instead of moving on, the model catches its own mistake before the next tool call: “That was wrong. Let me rewrite this cleanly with proper logic,” then writes a second, integer-cents-based version and, this time, calls run_pytest:

test_splitter.py::test_split_expense_three_people PASSED                 [ 25%]
test_splitter.py::test_split_expense_two_people PASSED                   [ 50%]
test_splitter.py::test_split_expense_requires_people PASSED              [ 75%]
test_splitter.py::test_split_expense_total_matches_original PASSED       [100%]

============================== 4 passed in 0.01s ==============================

All four tests pass. The agent’s own closing summary is worth reading, because it’s honest about something worth noticing: even after producing a correct fix and a passing test run, the model’s own prose narration of why the fix works briefly contradicts itself mid-sentence (“Wait, let me think again… Actually with my implementation…”) before landing on the right explanation. The code was correct. The model’s self-description of the code, in the same response, was momentarily wrong. That gap between “the agent says it’s right” and “the tests confirm it’s right” is exactly why Step 4 is verification, not narration.

Confirming the Fix Independently

Don’t take the agent’s own run_pytest call as the last word either. Re-run it yourself, and check the fix on an input the agent didn’t specifically target:

python -m pytest -v
python -c "
from splitter import split_expense
print(split_expense(10.00, ['Amir', 'Sam', 'Lee']))
print(split_expense(0.10, ['A', 'B', 'C']))
"
============================== 4 passed in 0.01s ==============================
{'Amir': 3.34, 'Sam': 3.33, 'Lee': 3.33}
{'A': 0.04, 'B': 0.03, 'C': 0.03}

Both sum exactly to the original amount, including the ten-cent case split three ways, an edge case nobody asked the agent to handle. The final, working splitter.py:

"""Split a shared expense evenly among a group of people."""


def split_expense(total, people):
    """Split `total` dollars evenly among `people` names.

    Returns a dict mapping each person's name to the amount they owe.
    """
    if not people:
        raise ValueError("Need at least one person to split an expense")

    num_people = len(people)

    # Work with cents (integers) to avoid floating point errors
    total_cents = int(round(total * 100))
    base_share_cents = total_cents // num_people

    # Calculate remainder in cents
    remaining_cents = total_cents % num_people

    result = {}
    for i, person in enumerate(people):
        share = base_share_cents + (1 if i < remaining_cents else 0)
        result[person] = round((share / 100.0), 2)

    return result

The fix works by never rounding a fraction of a cent in the first place: it converts the total to whole cents, divides with integer division to get an equal base share for everyone, then hands out the small number of leftover cents one at a time to the first few people. Nothing is ever rounded away, so the shares always sum back to the original total exactly.

Common Mistakes to Avoid

Trusting the Agent’s Summary Instead of the Test Output

This tutorial’s own transcript shows a model describing its own correct fix incorrectly, mid-response, before landing on the right explanation. Prose explanations are for your understanding; only a re-run of the actual test suite tells you whether the code is fixed.

Letting the Agent Weaken a Test to Make It Pass

Attempt 1’s suggested fix would have made the failing test pass without changing a single line of the buggy function. Whenever an agent’s diff touches the test file for a bug that’s supposed to live in the source file, read that diff first.

Not Forcing Verification Between Edits

The single largest difference between this tutorial’s broken run and its working run was one sentence added to the system prompt, requiring a run_pytest call after every write_file call. Without it, the agent wrote seven increasingly broken versions of the same function and never once checked any of them.

Giving a Bare Error Message Instead of a Reproducible Failure

“assert 9.99 == 10.0” and “a $10 dinner bill split three ways among Amir, Sam, and Lee comes up a cent short” describe the same bug, but only the second one, backed by a failing test and the actual source file, gave the agent enough to diagnose it correctly instead of guessing.

How to Confirm It All Works End to End

From a clean copy of the buggy splitter.py, the full workflow this tutorial taught looks like this: run the existing suite and see three tests pass; reproduce the reported bug by hand; write a new test that fails for exactly that reason; run python -m pytest -v and confirm one failure with a real traceback; run the tool-calling agent with a prompt that requires a run_pytest call after every edit; and finish by independently re-running the full suite yourself, plus at least one input the agent never saw. If all four tests pass on your own, unprompted re-run, and the diff only touched splitter.py, the bug is actually fixed.

Next Steps

  • Point this same loop at a real failing test from your own project instead of the expense splitter, and see whether the “read source, then test, then verify after every write” discipline holds up on unfamiliar code.
  • Swap the bespoke Ollama script for a hosted coding agent like Claude Code, Cursor, or Google’s Antigravity CLI. The methodology transfers directly: give it a failing test, give it real tools, and require it to verify before it stops, regardless of which agent is doing the editing.
  • If the code you’re debugging has no tests at all yet, start with sxz.io’s guide to writing characterization tests for untested legacy code before trying to fix anything, so you have a safety net to verify against in the first place.
  • For giving an agent durable context across sessions instead of re-explaining your codebase every time, see sxz.io’s guide to writing and organizing CLAUDE.md files.
  • Add more tools to the agent: a git diff tool to see what recently changed, or a grep tool to find every other caller of a buggy function, so the agent’s context-gathering isn’t limited to files you name for it.

Sources: pytest’s official documentation, including its pytest.approx() API reference, and Ollama’s official API documentation for the chat and tool-calling request format. Workflow inspired by Real Python’s “How to Debug Python Code With an AI Agent”, rebuilt here with an original bug, an original codebase, and a local Ollama tool-calling agent in place of Real Python’s Antigravity CLI walkthrough.

Featured image: U.S. Navy photograph of an Aviation Electronics Technician using a magnifier for detail work on a printed circuit board at Naval Air Station Jacksonville, Florida (PH2 D. Vukovich, 1987, National Archives at College Park, Still Pictures), a work of a U.S. military employee taken as part of official duties and in the public domain, via Wikimedia Commons; cropped, resized, and converted to WebP.

Tags:

AI AgentsDebuggingOllamapytestPython

Share

A blacksmith uses an industrial power hammer to forge glowing hot steel, illustrating Docker's push to harden software down to the OS package level
Previous Post

Docker’s Zero-CVE Push Turns Supply-Chain Security Into an OS-Level Default

Two U.S. Marines review a monitor together in a Network and Security Operations Center at Marine Corps Base Camp Pendleton
Next Post

GitLab Ships an Emergency Patch for a Critical Unauthenticated Code Injection Flaw

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
08 Oct
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
08 Oct
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
Trending
October 8, 2026
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
October 8, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
October 8, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google

Related Posts

A laptop wrapped in a chain and padlock, illustrating least-privilege controls for AI agents.
Learning Hub

How to Secure Tool-Using AI Agents Before They Touch Production

June 8, 2026
Colorful sticky notes arranged on an office wall, symbolizing governance checklists and planning.
Learning Hub

AI Governance for Agentic Apps: A Practical Checklist for Builders

June 8, 2026
A technician connects green fiber optic cables at a data center, representing a private production inference endpoint.
Learning Hub

How to Deploy a Fine-Tuned LLM Behind a Private Production Inference Endpoint

June 8, 2026
Narrow aisle behind black supercomputer racks in a data center
Learning Hub

Kubernetes SELinux Volume Labeling: What Cluster Operators Should Audit Before v1.37

June 8, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026