TRENDING
Rows of identical brass-colored apartment mailboxes with small locks and name labels along an orange corridor wall
October 9, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
Street-level upward view of the Monetary Authority of Singapore building and neighbouring office towers under a pale sky
October 9, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
Cast-iron late Qing dynasty coin minting press with a large flywheel, displayed in a museum case
October 9, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google
Rows of closed oak library card catalog drawers, each with a brass pull and a blank label holder
October 9, 2026
How to Encrypt PII in Python and Keep It Searchable With Blind Indexes
Close-up of a vintage Western Electric manual telephone switchboard with orange lamps, red patch cords plugged into jacks, a rotary dial and a black handset
October 9, 2026
Microsoft’s Agent Lightning v1.0 Turns Agent Training Into a Sample-Accounting Problem
09 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Two orange safety relief valves on grey pressure vessels in an industrial plant
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
Yellow diamond-shaped merging traffic warning sign showing a side road joining a main road
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A lugworm lying on wet sand and mud at low tide
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 232 Posts
News 234 Posts
Learning Hub 204 Posts
Home/Learning Hub/How to Build a Production-Grade AI Agent With Pydantic AI and Ollama
Learning Hub

How to Build a Production-Grade AI Agent With Pydantic AI and Ollama

A step-by-step guide to building a local AI agent with Pydantic AI: structured output, tool calling, dependency injection, retries, and tests that never call a real model.

August 14, 2026 15 Min Read
62

If you have ever built an AI agent by calling an LLM SDK directly, you already know the pattern: the prototype works in a notebook, then you move it toward production and start bolting on patches. A try/except around json.loads. A helper that strips markdown fences off the model’s reply. A hand-written JSON schema for every tool. A retry loop you copy-paste into every new agent. None of these are hard by themselves, but together they become most of your codebase, and the actual agent logic disappears underneath the glue.

Table Of Content

  • What You Will Build
  • Prerequisites
  • Step 1: Set Up Your Environment
  • Step 2: Get Real Structured Output From Your First Agent
  • The problem this solves
  • Build the model and the agent
  • Run it
  • What is actually happening under the hood
  • Step 3: Give the Agent a Tool It Can Call
  • The problem this solves
  • Register the tool
  • Watch the agent decide to call it
  • Step 4: Pass Runtime Context With Dependency Injection
  • The problem this solves
  • Define a deps type
  • Compare two runs side by side
  • Step 5: Make Bad Tool Input a Recoverable Error, Not a Crash
  • What ModelRetry does
  • An honest gotcha: a capable model may just avoid your test scenario
  • Step 6: Test the Agent Without Ever Calling a Real Model
  • Why this matters
  • Two fake models, two different jobs
  • TestModel: a fast structural smoke test
  • FunctionModel: scripting an exact scenario
  • Run the tests
  • Step 7: Swap Models With a One-Line Change, Carefully
  • The problem this solves
  • Point the same agent at a different local model
  • An honest gotcha: smaller models are not drop-in interchangeable
  • A working fallback for weaker models
  • Verify Everything Works End to End
  • Common Mistakes and Gotchas (Recap)
  • Next Steps

Pydantic AI is a Python agent framework, built by the team behind the Pydantic validation library, that replaces that glue with a small number of well-tested primitives: typed structured output, auto-generated tool schemas, dependency injection for runtime context, built-in retry handling, and a model-agnostic interface so the same agent code runs against OpenAI, Anthropic, or a model running on your own machine.

This tutorial builds a real IT support ticket triage agent from scratch, entirely on your own computer with Ollama, so there is no API key and no cloud bill. Every command and every output shown below was run for real while writing this post, including two genuine surprises: a testing tool that broke against our own validation logic, and a smaller model that failed outright under the default settings. Both are explained, and both are fixed, because that is exactly the kind of thing you will hit the first time you try this yourself.

What You Will Build

An agent that reads a raw support ticket (the kind of free-text message a user might type into a help desk form) and returns a typed TicketTriage object: a category, a severity, a confidence score, and a plain-English summary. Along the way the agent will:

  • Call a tool to check whether a known service outage explains the ticket
  • Use per-request context (a customer’s support tier) to change its own behavior, without a single global variable
  • Reject bad input from itself and try again, instead of silently returning nonsense
  • Run through a full test suite that never once calls a real language model
  • Switch to a different local model with a one-line change, with the caveats that come with it

Prerequisites

  • A computer that can run Python 3.10 or newer (this tutorial used Python 3.13.14 on Windows, but the code is not Windows-specific)
  • Comfort with basic Python: functions, type hints, and what a dataclass is
  • Ollama installed and running locally, with a tool-capable model already pulled. This tutorial uses qwen3.5:4b (about 3.4 GB on disk); pull it with ollama pull qwen3.5:4b if you do not already have it
  • No API keys, no cloud account, and no cost: everything below talks to Ollama on localhost

Confirm Ollama is actually running and reachable before you start:

curl http://localhost:11434/api/version

You should get back something like {"version":"0.32.9"}. If that call fails, start Ollama first (on most installs, running the ollama app or ollama serve from a terminal) and confirm your model is pulled with ollama list before moving on.

Step 1: Set Up Your Environment

Create an isolated virtual environment and install Pydantic AI. The plain pydantic-ai package (as opposed to the smaller pydantic-ai-slim it depends on) pulls in the openai Python package too, which is what Pydantic AI’s Ollama support is built on: Ollama exposes an OpenAI-compatible chat API, so Pydantic AI talks to it the same way it talks to OpenAI itself.

python -m venv venv
venv\Scripts\activate   # on Linux/macOS: source venv/bin/activate
pip install pydantic-ai pytest

This tutorial was written against pydantic-ai 2.30.0 and pytest 9.1.1. Confirm your install:

python -c "import pydantic_ai; print(pydantic_ai.__version__)"

Step 2: Get Real Structured Output From Your First Agent

The problem this solves

If you call a chat completion API directly and ask the model to “reply in JSON,” you get back a string. Maybe it is valid JSON. Maybe it is valid JSON wrapped in a markdown code fence. Maybe a field is missing, or a number came back as the string "high" instead of an actual number. You find out at runtime, usually in production, usually from a stack trace three layers deep in your own parsing code.

Pydantic AI’s fix is to let you hand the agent a real Pydantic model as its output_type. The agent guarantees that result.output is either a validated instance of that exact class, or the run raises a clear exception. There is no string to parse.

Build the model and the agent

Create triage_agent.py:

from pydantic import BaseModel, Field
from typing import Literal
from pydantic_ai import Agent
from pydantic_ai.models.ollama import OllamaModel
from pydantic_ai.providers.ollama import OllamaProvider


class TicketTriage(BaseModel):
    category: Literal["network", "hardware", "software", "account", "security", "other"]
    severity: Literal["low", "medium", "high", "critical"]
    confidence: float = Field(ge=0.0, le=1.0)
    summary: str = Field(description="One-sentence plain-English summary of the issue")


model = OllamaModel(
    "qwen3.5:4b",
    provider=OllamaProvider(base_url="http://localhost:11434/v1"),
)

triage_agent = Agent(
    model,
    output_type=TicketTriage,
    system_prompt=(
        "You are an IT support ticket triage assistant. "
        "Read the raw ticket text and classify it."
    ),
)

if __name__ == "__main__":
    ticket = (
        "Subject: Cannot reach the internal wiki\n"
        "Body: Since about 9am I get a 'connection timed out' error when I try to "
        "open wiki.corp.internal from the office. VPN is connected. Three other "
        "people on my team report the same thing. We use this constantly for "
        "runbooks so this is blocking us."
    )
    result = triage_agent.run_sync(ticket)
    print(result.output.model_dump_json(indent=2))

A few things worth noticing before you run it: Literal[...] restricts category and severity to an exact fixed set of strings, and Field(ge=0.0, le=1.0) puts a real numeric bound on confidence. Both of these are ordinary Pydantic validation, and Pydantic AI enforces them on whatever the model produces, not just on data your own code constructs.

Run it

python triage_agent.py

Real output from this exact script, unedited:

{
  "category": "network",
  "severity": "high",
  "confidence": 0.95,
  "summary": "Multiple team members experiencing connection timeouts when attempting to access the internal wiki service from the office, indicating a potential network or service availability issue affecting multiple users simultaneously."
}

result.output is a real TicketTriage instance here, not a dict and not a string. You can access result.output.severity directly and your type checker knows it is one of the four literal strings you defined.

What is actually happening under the hood

Here is something that is not obvious from the code above, and that you will only see if you inspect the full message history with result.all_messages(): even with zero tools registered, Pydantic AI’s default output mode implements structured output as a synthetic tool call named final_result. Printing the tool-related message parts for the run above shows exactly this:

ToolCallPart final_result {"category":"network","severity":"high","confidence":0.95,"summary":"Multiple team members experiencing..."}
ToolReturnPart final_result Final result processed.

In other words, Pydantic AI does not ask the model to “please reply in JSON” and hope. It tells the model there is a tool called final_result whose parameters are your Pydantic model’s schema, and the model has to call it correctly to finish the run. This is why the validation is reliable: it is enforced the same way tool-calling itself is enforced, not bolted on afterward. Keep this in mind, it becomes directly relevant in Step 7.

Step 3: Give the Agent a Tool It Can Call

The problem this solves

Defining a tool for a raw LLM SDK usually means hand-writing a JSON Schema dictionary that mirrors your Python function’s signature, then keeping the two in sync by hand every time you change either one. It is boilerplate that actively invites drift.

Pydantic AI generates the schema for you from the function’s type hints and docstring. You write one normal Python function, and that is the only source of truth.

Register the tool

Add this to your script, and add related_outage: bool to TicketTriage:

@triage_agent.tool_plain
def check_service_status(service_name: str) -> str:
    """Look up the current status of an internal service.

    Args:
        service_name: The internal service name, e.g. 'wiki', 'vpn', 'email', 'sso'.
    """
    status_board = {
        "wiki": "outage: intermittent 502s since 08:47 UTC, engineering investigating",
        "vpn": "operational",
        "email": "operational",
        "sso": "degraded: slow logins reported in the last 30 minutes",
    }
    return status_board.get(service_name.lower(), "no status entry found for that service")

@tool_plain is for tools that need no access to per-run context. (Step 4 introduces @tool, the version that does.) The docstring’s Args: section is not decoration, Pydantic AI parses it to fill in each parameter’s description in the generated schema, so the model sees the same explanation you wrote for other developers.

Watch the agent decide to call it

Run the same wiki ticket through the updated agent and print the full tool-call trace:

{
  "category": "network",
  "severity": "high",
  "confidence": 0.95,
  "summary": "Multiple users experiencing connection timeout when accessing internal wiki due to ongoing 502 server errors, confirmed by engineering outage status since early morning hours",
  "related_outage": true
}

--- trace ---
ToolCallPart check_service_status {"service_name":"wiki"}
ToolReturnPart outage: intermittent 502s since 08:47 UTC, engineering investigating
ToolCallPart final_result {"category":"network","severity":"high","confidence":0.95,"summary":"...","related_outage":true}
ToolReturnPart Final result processed.

Nobody told the agent when to call check_service_status. It read the ticket, decided a service-status lookup was relevant, extracted "wiki" as the argument on its own, used the real answer to set related_outage: true, and only then called final_result. That decision-making is the actual value of a tool-using agent; the schema generation is just what makes it safe to build one quickly.

Step 4: Pass Runtime Context With Dependency Injection

The problem this solves

Real tools usually need something that is not part of the conversation: a database connection, the current user’s ID, a feature flag, or, in this tutorial, which support tier a customer is on. The tempting shortcuts are a module-level global variable (breaks the moment you handle two requests at once) or a closure over some outer variable (works, but scatters state in a way that is hard to test). Pydantic AI has a third option built in.

Define a deps type

from dataclasses import dataclass

@dataclass
class TriageDeps:
    """Runtime context for one triage run. No globals, no closures."""
    customer_tier: Literal["standard", "vip"]
    status_board: dict[str, str]

Tell the agent about it with deps_type, and change the tool to accept a RunContext[TriageDeps] as its first parameter:

from pydantic_ai import RunContext

triage_agent = Agent(
    model,
    deps_type=TriageDeps,
    output_type=TicketTriage,
    system_prompt=(
        "You are an IT support ticket triage assistant. Classify the ticket and "
        "use check_service_status to see if a known outage explains it. "
        "VIP customers should never be scored below 'high' severity even for "
        "minor issues, because their contracts guarantee faster response."
    ),
)

@triage_agent.tool
def check_service_status(ctx: RunContext[TriageDeps], service_name: str) -> str:
    """Look up the current status of an internal service.

    Args:
        service_name: The internal service name, e.g. 'wiki', 'vpn', 'email', 'sso'.
    """
    board = ctx.deps.status_board
    return board.get(service_name.lower(), "no status entry found for that service")

@triage_agent.system_prompt
def add_customer_tier(ctx: RunContext[TriageDeps]) -> str:
    return f"This customer's support tier is: {ctx.deps.customer_tier}."

ctx.deps is whatever you pass to run_sync(..., deps=...). The tool reads it directly, with no global state anywhere. The @triage_agent.system_prompt function is a second, dynamic way to shape the model’s instructions per run: it re-runs on every call and its return value gets appended to the static system_prompt string, so the model finds out the customer’s tier fresh each time instead of it being baked into a prompt template.

Compare two runs side by side

Run the identical ticket (a mildly annoying but non-urgent slow VPN) through both a standard and a VIP customer:

=== standard customer ===
{
  "category": "network",
  "severity": "medium",
  "confidence": 0.65,
  "summary": "VPN connection experiencing elevated latency approximately 2x normal speed, but service remains fully operational and usable for current needs...",
  "related_outage": false
}

=== VIP customer, same ticket ===
{
  "category": "software",
  "severity": "high",
  "confidence": 0.65,
  "summary": "VIP customer reporting 2x normal latency on VPN connection... VIP tier requires minimum high severity rating regardless of perceived criticality per contractual response time guarantees.",
  "related_outage": false
}

Same ticket text, same tool, same code path, different runtime context, different outcome. That is dependency injection working as intended. One honest note: the category field also changed between the two runs (network versus software), which is not something the tier logic asked for. That is ordinary small-model non-determinism on a genuinely ambiguous ticket (a VPN issue could reasonably be filed under either), not a bug in the dependency injection itself. If you need identical categorization across repeated runs, set a fixed seed and low temperature in your model settings, and expect a genuinely ambiguous ticket to still wobble occasionally even then.

Step 5: Make Bad Tool Input a Recoverable Error, Not a Crash

What ModelRetry does

Sometimes the model calls your tool with an argument that does not make sense, a misspelled service name, an ID that does not exist, a date in the wrong format. Raising a normal Python exception inside a tool ends the entire run. Pydantic AI gives you a better option: raise pydantic_ai.ModelRetry with a helpful message, and the framework feeds that message back to the model as the tool’s result, letting it try again instead of crashing.

from pydantic_ai import ModelRetry

@triage_agent.tool
def check_service_status(ctx: RunContext[TriageDeps], service_name: str) -> str:
    """Look up the current status of an internal service.

    Args:
        service_name: The exact internal service key: one of 'wiki', 'vpn', 'email', 'sso'.
    """
    board = ctx.deps.status_board
    key = service_name.lower().strip()
    if key not in board:
        raise ModelRetry(
            f"'{service_name}' is not a valid service key. "
            f"Valid keys are exactly: {', '.join(sorted(board))}."
        )
    return board[key]

Give the agent a retry budget with Agent(..., retries=3), since the default is a single attempt.

An honest gotcha: a capable model may just avoid your test scenario

To see the retry path fire, the plan was to send a ticket about “the ticketing system” and watch the model guess an invalid key like ticketing. It did not cooperate. Across several attempts, including a ticket that used the literal word “ticketing” three times, qwen3.5:4b either mapped it correctly to sso on the first try, or checked a couple of valid keys (email, then sso) before answering, never once passing an invalid key to the tool:

CALL check_service_status {"service_name":"email"}
RETURN operational
CALL check_service_status {"service_name":"sso"}
RETURN degraded: slow logins reported in the last 30 minutes
CALL final_result {"category":"software","severity":"medium",...}

That is good news for production (a reasonably capable local model naturally avoids the mistake you were worried about) and bad news for testing (you cannot reliably reproduce that failure path just by phrasing a prompt cleverly, because you do not control what a live model decides to do on a given run). That is exactly the gap Step 6 closes.

Step 6: Test the Agent Without Ever Calling a Real Model

Why this matters

A test suite that calls a real LLM on every run is slow, costs money (or, with a local model, competes for your GPU), and can flake for reasons that have nothing to do with your code. Pydantic AI ships two fake “models” built specifically for testing, and they solve two different problems.

Two fake models, two different jobs

TestModel: a fast structural smoke test

TestModel auto-generates a schema-valid response without any network call. It does not prove your agent is smart, it proves the plumbing is not broken: the tools are registered correctly, the output schema is satisfiable, nothing has a typo. Use agent.override(model=...) to swap the model for the duration of a test:

from pydantic_ai.models.test import TestModel

def test_agent_shape_with_test_model():
    agent = build_agent()
    with agent.override(model=TestModel(call_tools=[])):
        result = agent.run_sync("wiki is down for everyone", deps=TEST_DEPS)
    assert isinstance(result.output, TicketTriage)

That call_tools=[] is not decoration, it is a fix for a real bug this tutorial ran into. TestModel‘s default behavior is to call every registered tool once with auto-generated placeholder arguments, things like service_name="a". Our check_service_status tool validates its input and immediately raises ModelRetry for anything not in the status board, “a” included. TestModel does not read or learn from a ModelRetry message, it just submits another placeholder guess, so the run failed after exhausting its retries with UnexpectedModelBehavior: Tool 'check_service_status' exceeded max retries count of 3. The fix is one keyword argument: tell TestModel not to call that tool at all when the test only cares about the output shape.

FunctionModel: scripting an exact scenario

Since Step 5 showed a real model dodging the bad-input path, FunctionModel is how you force it deterministically. You write a plain Python function that plays the role of the model: it receives the message history and returns whatever response you want, call by call.

from pydantic_ai.messages import ModelResponse, RetryPromptPart, ToolCallPart
from pydantic_ai.models.function import AgentInfo, FunctionModel

def test_retry_path_with_function_model():
    call_count = {"n": 0}

    def scripted_model(messages, info: AgentInfo) -> ModelResponse:
        call_count["n"] += 1
        if call_count["n"] == 1:
            # Deliberately use an invalid service key.
            return ModelResponse(parts=[
                ToolCallPart("check_service_status", {"service_name": "ticketing"})
            ])
        if call_count["n"] == 2:
            # Confirm we actually got a RetryPromptPart, then "notice the hint".
            retry_parts = [p for p in messages[-1].parts if isinstance(p, RetryPromptPart)]
            assert retry_parts and "not a valid service key" in str(retry_parts[0].content)
            return ModelResponse(parts=[
                ToolCallPart("check_service_status", {"service_name": "sso"})
            ])
        return ModelResponse(parts=[
            ToolCallPart("final_result", {
                "category": "software", "severity": "medium", "confidence": 0.8,
                "summary": "SSO degradation likely affecting ticketing portal logins.",
                "related_outage": True,
            })
        ])

    agent = build_agent()
    with agent.override(model=FunctionModel(scripted_model)):
        result = agent.run_sync("the ticketing portal is down", deps=TEST_DEPS)

    assert call_count["n"] == 3
    assert result.output.related_outage is True

Run the tests

pytest test_agent.py -v -s
test_agent.py::test_agent_shape_with_test_model PASSED
test_agent.py::test_retry_path_with_function_model
Retry path exercised correctly: 3 model calls, final output = {"category":"software","severity":"medium","confidence":0.8,"summary":"SSO degradation likely affecting ticketing portal logins.","related_outage":true}
PASSED

2 passed, 1 warning in 0.72s

(The one warning is an internal DeprecationWarning from a library dependency about event-loop handling, unrelated to anything in this tutorial’s code.) Both tests run in well under a second, with zero network calls and zero dependency on what a live model feels like doing that day. The ModelRetry path from Step 5, the one the real model refused to trigger organically, is now covered by an assertion that actually fires every single time.

Step 7: Swap Models With a One-Line Change, Carefully

The problem this solves

With a raw SDK, moving from one provider to another (or from a hosted model to a local one) usually means rewriting your request-building and response-parsing code, because every provider shapes tool calls and messages differently. A Pydantic AI agent is written against the framework’s own message types, so the same agent code runs against a different model by changing what you pass in as the model argument.

Point the same agent at a different local model

from pydantic_ai.models.ollama import OllamaModel
from pydantic_ai.providers.ollama import OllamaProvider

m = OllamaModel("qwen2.5:1.5b", provider=OllamaProvider(base_url="http://localhost:11434/v1"))
result = triage_agent.run_sync(ticket, deps=deps, model=m)

That is the entire change: one string, from qwen3.5:4b to the much smaller qwen2.5:1.5b (986 MB versus 3.4 GB). No tool code changed, no output schema changed.

An honest gotcha: smaller models are not drop-in interchangeable

Running the exact email-outage ticket from Step 4 against qwen2.5:1.5b, even with retries=3, did not produce a passing result. It failed consistently:

FAILED after 7.5 s: UnexpectedModelBehavior Exceeded maximum output retries (3)

Ollama reports qwen2.5:1.5b as supporting the tools capability, and it does accept tool-calling requests, but at 1.5 billion parameters it could not reliably complete the exact final_result tool call (recall from Step 2: structured output is itself a tool call under Pydantic AI’s default mode) closely enough to pass schema validation, even across three attempts. The one-line model swap really is one line of code, but code compatibility and model capability are two different things, and only one of them is guaranteed by the framework.

A working fallback for weaker models

Pydantic AI has more than one way to get structured output, and the default tool-calling mode (ToolOutput) is not always the best fit for a small model. Switching to PromptedOutput, which asks the model to produce the matching JSON directly in its text reply instead of via a tool call, fixed it:

from pydantic_ai.output import PromptedOutput

agent = Agent(
    "test",
    deps_type=TriageDeps,
    output_type=PromptedOutput(TicketTriage),  # was: output_type=TicketTriage
    retries=3,
    system_prompt="...",
)
SUCCESS in 1.6 s
{
  "category": "software",
  "severity": "critical",
  "confidence": 0.9,
  "summary": "The email service is currently experiencing outages.",
  "related_outage": true
}

Same small model, same ticket, one output-mode change, and it succeeded in a fraction of the time the failing attempts took. The lesson is not “always use PromptedOutput”. It is that when you swap in a smaller or less capable model, check whether it is actually completing runs successfully before you trust it, and know that the output mode is a second lever you can pull if the default is not working.

Verify Everything Works End to End

  1. Confirm Ollama is running: curl http://localhost:11434/api/version returns a version string.
  2. Confirm your model is pulled: ollama list shows it.
  3. Run the Step 4 script directly and confirm you get back a TicketTriage object with a populated related_outage field, not an exception.
  4. Run pytest test_agent.py -v -s and confirm both tests pass in well under a second. If test_agent_shape_with_test_model fails with UnexpectedModelBehavior, you almost certainly hit the TestModel plus tool-validation gotcha from Step 6, add call_tools=[].
  5. If you try Step 7’s model swap yourself, do not assume success, check the actual output. A model can be listed as tool-capable by Ollama and still fail to complete a structured-output run.

Common Mistakes and Gotchas (Recap)

  • Assuming TestModel is a full behavioral test. It only proves your schema and tool registration are not broken. It does not, and cannot, tell you whether the agent’s actual decisions are good.
  • TestModel versus tools with real validation. TestModel calls every tool with placeholder arguments by default. If a tool validates its input and raises ModelRetry on nonsense, TestModel will not learn from that message, it just keeps guessing, and the run can fail. Use call_tools=[] when you only need to test output shape.
  • Forgetting to raise the retry budget. Agent(...) defaults to a small number of attempts. A tool that validates its own input needs enough retries to actually recover from a bad first guess.
  • Treating “the code runs” as “the model works.” Swapping model= is genuinely one line. Whether the new model can reliably complete your agent’s task is a separate question you have to check empirically, every time.

Next Steps

From here, reasonable next steps are: add a second tool with its own ModelRetry validation and a test that scripts its failure path with FunctionModel; try NativeOutput as a third structured-output mode and compare its behavior on your smaller model against PromptedOutput; and, if you eventually point this agent at a hosted provider instead of Ollama, notice that everything from Step 2 through Step 6 (the model, the tools, the tests) does not change, only the model= line does, which is the entire point of building on a framework instead of a provider’s raw SDK.

Tags:

AI AgentsOllamaPydantic AIPythonSoftware Testing

Share

The C-Lion1 submarine fiber optic cable landing station building in Rostock-Markgrafenheide, Germany
Previous Post

DigitalOcean’s Data Locality Tax Turns RAG Latency Into a Geography Problem

Macro photo of a single marbled six-sided die with silver pips, representing the randomness Anthropic's watermark quietly steers
Next Post

Anthropic to Watermark Claude’s AI-Generated Text to Comply With the EU AI Act

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
08 Oct
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
08 Oct
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
Trending
October 8, 2026
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
October 8, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
October 8, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google

Related Posts

A laptop wrapped in a chain and padlock, illustrating least-privilege controls for AI agents.
Learning Hub

How to Secure Tool-Using AI Agents Before They Touch Production

June 8, 2026
Colorful sticky notes arranged on an office wall, symbolizing governance checklists and planning.
Learning Hub

AI Governance for Agentic Apps: A Practical Checklist for Builders

June 8, 2026
A technician connects green fiber optic cables at a data center, representing a private production inference endpoint.
Learning Hub

How to Deploy a Fine-Tuned LLM Behind a Private Production Inference Endpoint

June 8, 2026
Narrow aisle behind black supercomputer racks in a data center
Learning Hub

Kubernetes SELinux Volume Labeling: What Cluster Operators Should Audit Before v1.37

June 8, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026