TRENDING
Rows of identical brass-colored apartment mailboxes with small locks and name labels along an orange corridor wall
October 9, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
Street-level upward view of the Monetary Authority of Singapore building and neighbouring office towers under a pale sky
October 9, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
Cast-iron late Qing dynasty coin minting press with a large flywheel, displayed in a museum case
October 9, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google
Rows of closed oak library card catalog drawers, each with a brass pull and a blank label holder
October 9, 2026
How to Encrypt PII in Python and Keep It Searchable With Blind Indexes
Close-up of a vintage Western Electric manual telephone switchboard with orange lamps, red patch cords plugged into jacks, a rotary dial and a black handset
October 9, 2026
Microsoft’s Agent Lightning v1.0 Turns Agent Training Into a Sample-Accounting Problem
09 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Two orange safety relief valves on grey pressure vessels in an industrial plant
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
Yellow diamond-shaped merging traffic warning sign showing a side road joining a main road
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A lugworm lying on wet sand and mud at low tide
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 232 Posts
News 234 Posts
Learning Hub 204 Posts
Home/Learning Hub/How to Get Reliable Structured JSON Output From Ollama Models in Python
Learning Hub

How to Get Reliable Structured JSON Output From Ollama Models in Python

Ollama can constrain a local model's output to a JSON schema instead of hoping it replies in valid JSON; here is how that works, what it actually enforces, and two real bugs found while testing it.

August 16, 2026 16 Min Read
70

If you have ever asked a local LLM to “reply in JSON” and then wrapped the result in a try/except block, you already know the problem this tutorial solves. The model is generating text, one token at a time, based on what sounds statistically plausible. It has no built-in concept of “this field is required” or “this must be a boolean.” Most of the time it gets close enough. Occasionally it wraps the JSON in a markdown code fence, renames a field, or drifts into a completely different shape than you asked for, and your parser breaks in production.

Table Of Content

  • What You Will Build
  • Prerequisites
  • Step 1: See Why “Just Ask for JSON” Fails
  • The problem
  • Run it
  • Step 2: Try Unconstrained JSON Mode
  • The problem this solves, partially
  • Run it three times in a row
  • Step 3: Force a Real Schema With Structured Outputs
  • Define the shape with Pydantic
  • Run it three times
  • Step 4: Build a Realistic Extraction Task
  • Run it against three real tickets
  • Step 5: Know What the Schema Actually Enforces
  • Smaller models stay valid, but get sloppier
  • Array length and numeric range constraints are enforced
  • Regex patterns are not enforced, they fail the request outright
  • Step 6: The Reasoning Model Gotcha
  • The fix
  • Step 7: Wrap It in a Retry Helper for Production
  • Watch it succeed
  • Watch it fail honestly, and see why retries are not a fix for everything
  • Common Mistakes and How to Verify Your Setup
  • Next Steps

Ollama has a real fix for this called structured outputs: you hand it a JSON Schema (a formal, machine-readable description of the exact shape your data must have), and Ollama constrains the model’s token generation so the output is guaranteed to match that shape. This is not prompt engineering. Ollama compiles the schema into a generation grammar and applies it during sampling, so the model cannot emit a token that would violate the schema’s structure. You can see this mechanism directly later in this tutorial: a schema Ollama cannot compile fails the request outright with the error “failed to parse grammar,” rather than the model simply ignoring it.

This tutorial builds that up from first principles, entirely on your own machine with Ollama and Python, no API key and no cloud bill. Every command and every piece of output below was run for real while writing this post, including two genuine bugs discovered along the way: a reasoning model that silently returns nothing under the default settings, and a schema constraint that Ollama accepts on paper but actually rejects at request time. Both are reproduced, explained, and fixed.

What You Will Build

A small Python library function, extract_structured(), that takes a prompt and a Pydantic model and returns a validated instance of that model, never a raw string you have to parse yourself. Along the way you will:

  • Reproduce, with real captured output, exactly how “just ask for JSON” fails
  • Use Ollama’s format parameter to constrain a model to a JSON Schema
  • Build a realistic support-ticket triage extractor with enums and list fields
  • Find out empirically which JSON Schema keywords Ollama actually enforces, and which ones cause a request to fail outright
  • Reproduce and fix a real bug where a reasoning model returns an empty response under structured output
  • Wrap all of it in a retry helper suitable for production use

Prerequisites

  • A computer that can run Python 3.10 or newer (this tutorial used Python 3.13.14 on Windows, but nothing here is Windows-specific)
  • Basic Python comfort: functions, type hints, and what a class is
  • Ollama installed and running locally. Structured outputs need Ollama 0.5.0 or newer; this tutorial used 0.32.9. Check yours with ollama --version
  • At least one model pulled. This tutorial mainly uses qwen2.5:1.5b (about 1 GB), a small, fast model that is a good default for structured extraction tasks. Pull it with ollama pull qwen2.5:1.5b
  • No API keys, no cloud account, no cost: everything below talks to Ollama on localhost

Confirm Ollama is running and reachable before you start:

curl http://localhost:11434/api/version

You should get back something like {"version":"0.32.9"}. If that fails, start Ollama first (run the ollama app, or ollama serve from a terminal) before continuing.

Create an isolated environment and install the two libraries this tutorial needs:

python -m venv venv
venv\Scripts\activate   # on Linux/macOS: source venv/bin/activate
pip install requests pydantic

This was written against requests 2.34.2 and pydantic 2.13.4. Everything here talks to Ollama’s raw HTTP API directly with requests, on purpose: understanding the actual request and response shape makes it much easier to debug things later, whichever framework or SDK you eventually build on top of it.

Step 1: See Why “Just Ask for JSON” Fails

The problem

Start with the most natural thing to try: describe the fields you want in plain English and hope for the best. Save this as step1_naive.py:

import json
import requests

OLLAMA_URL = "http://localhost:11434/api/generate"
MODEL = "qwen2.5:1.5b"

review = "This blender is amazing, it crushed ice in seconds! Shipping took 2 weeks though, kind of annoying. 5 stars overall."

prompt = f"""Extract the product rating (1-5), shipping complaint (true/false), and a one-sentence summary from this review as JSON:

"{review}"
"""

resp = requests.post(OLLAMA_URL, json={
    "model": MODEL,
    "prompt": prompt,
    "stream": False,
})
resp.raise_for_status()
raw_text = resp.json()["response"]
print("=== RAW MODEL OUTPUT ===")
print(raw_text)
print("=== ATTEMPTING json.loads() ===")
try:
    parsed = json.loads(raw_text)
    print("SUCCESS:", parsed)
except json.JSONDecodeError as e:
    print("FAILED:", e)

Run it

python step1_naive.py

Real output from this exact script, unedited:

=== RAW MODEL OUTPUT ===
```json
{
  "rating": {
    "product_rating": 5,
    "shipping_complaint": false
  },
  "summary": "This blender is exceptional at crushing ice quickly; however, the shipping took longer than expected and could be considered annoying. Overall, it gets 5 stars."
}
```
=== ATTEMPTING json.loads() ===
FAILED: Expecting value: line 1 column 1 (char 0)

Two separate problems in that one response. First, the model wrapped its answer in a markdown code fence (the triple backticks), so json.loads() chokes on the backtick character before it even reaches the JSON. Second, even if you stripped the fence, the shape is wrong: you asked for a flat rating field and got a nested object with a differently-named product_rating field inside it instead. A parser built around the shape you asked for would break on both counts.

Step 2: Try Unconstrained JSON Mode

The problem this solves, partially

Ollama has a simpler, older feature for this: set "format": "json" and it will structure the response as a syntactically valid JSON object. Note the wording: valid JSON, not JSON matching any particular shape. Change the request in step1_naive.py to add the format field, and make sure your prompt mentions JSON explicitly (Ollama’s own docs warn that skipping this can make the model generate large amounts of whitespace):

resp = requests.post(OLLAMA_URL, json={
    "model": MODEL,
    "prompt": prompt + "\nRespond using JSON.",
    "format": "json",
    "stream": False,
})

Run it three times in a row

Real, unedited output from three separate runs of the identical request:

--- run 1 ---
{
  "product_rating": 5,
  "shipping_complaint": false,
  "summary": "This blender is highly rated at 5 stars for its quick ice-crushing ability but had a shipping complaint due to the long wait time."
}

--- run 2 ---
{
  "rating": 5,
  "shipping_complaint": false,
  "summary": "This blender is great and fast at crushing ice, but shipping was delayed slightly for a minor inconvenience"
}

--- run 3 ---
{
  "rating": 5,
  "shipping_complaint": false,
  "summary": "This blender is great, but shipping was a bit slow, but overall it's a five-star product."
}

Every run produced valid JSON this time, real progress over Step 1. But look closely at the field name: run 1 called it product_rating, runs 2 and 3 called it rating. If your downstream code does data["rating"], it silently breaks about a third of the time, and you would not notice until something consumed a missing field as None. JSON mode guarantees syntax, not a contract. You still need something that pins down the actual field names and types.

One more thing worth noticing here, since it comes up again later: every single run above marked shipping_complaint as false, despite the review explicitly calling the two-week shipping time “kind of annoying.” Keep that in mind. It becomes relevant in Step 4.

Step 3: Force a Real Schema With Structured Outputs

Define the shape with Pydantic

Pydantic is a Python validation library: you describe your data as a class with typed fields, and Pydantic can both generate a JSON Schema from that class and validate incoming data against it. That combination is exactly what you need here. Instead of describing the fields in a prompt and hoping, define them as a real Python class:

from pydantic import BaseModel

class ReviewExtraction(BaseModel):
    rating: int
    shipping_complaint: bool
    summary: str

ReviewExtraction.model_json_schema() turns that class into a JSON Schema automatically. Pass that schema as Ollama’s format parameter instead of the string "json", and Ollama constrains generation to match it exactly:

import json
import requests
from pydantic import BaseModel, ValidationError

class ReviewExtraction(BaseModel):
    rating: int
    shipping_complaint: bool
    summary: str

resp = requests.post(OLLAMA_URL, json={
    "model": MODEL,
    "prompt": prompt,
    "format": ReviewExtraction.model_json_schema(),
    "stream": False,
})
resp.raise_for_status()
raw_text = resp.json()["response"]
print(raw_text)

extraction = ReviewExtraction.model_validate_json(raw_text)
print(extraction)

Run it three times

--- run 1 ---
{
    "rating": 5,
    "shipping_complaint": false,
    "summary": "This blender is fantastic for its ability to quickly crush ice and was surprisingly quick with delivery despite the delay in shipping."
}
rating=5 shipping_complaint=False summary='...'

--- run 2 ---
{
  "rating": 5,
  "shipping_complaint": false,
  "summary": "This blender is awesome for its performance but the shipping was a bit slow and inconvenient."
}
rating=5 shipping_complaint=False summary='...'

--- run 3 ---
{ "rating": 5, "shipping_complaint": false, "summary": "This blender is excellent for quick ice crushing, despite a slightly long shipping time." }
rating=5 shipping_complaint=False summary='...'

Every run used the exact field names from the Pydantic model, every time, and model_validate_json() passed without raising. That consistency is the actual feature: the model is not being asked nicely to use the right field names, it is mechanically prevented from doing anything else while it is generating the response.

Notice something else, though: shipping_complaint is still false in every run, exactly like in Step 2, even though nothing about adding the schema changed the model’s understanding of the review. This is the most important thing to internalize about structured outputs: the schema guarantees the shape of the answer, not the correctness of the answer. A model that is confidently wrong about the content will now be confidently wrong in a format your code can parse without crashing. Schema validation and fact-checking are two different problems, and you still need both.

Step 4: Build a Realistic Extraction Task

A single flat object with a boolean and a string is a good way to see the mechanism work, but real extraction tasks usually need enums (a fixed set of allowed string values) and list fields. Here is a support-ticket triage schema that uses both, saved as triage.py:

from enum import Enum
from typing import Literal
from pydantic import BaseModel

class Severity(str, Enum):
    low = "low"
    medium = "medium"
    high = "high"
    critical = "critical"

class TicketTriage(BaseModel):
    category: Literal["billing", "technical", "account", "other"]
    severity: Severity
    sentiment: Literal["positive", "neutral", "negative"]
    action_items: list[str]

Literal[...] and Enum both compile down to an enum array in the generated JSON Schema, so the model is constrained to exactly those strings, nothing else can appear in that field. list[str] compiles to a JSON Schema array of strings.

Run it against three real tickets

Send three genuinely different tickets through the same schema:

import time

TICKETS = [
    "I've been charged twice for my subscription this month and nobody has responded to my last two emails. This is unacceptable, I want a refund immediately.",
    "Quick question, how do I change the email address on my account? Not urgent, just whenever you get a chance.",
    "The app crashes every time I try to export a report. This is blocking our entire team from finishing quarterly numbers before the board meeting tomorrow morning.",
]

for ticket in TICKETS:
    start = time.time()
    resp = requests.post(OLLAMA_URL, json={
        "model": "qwen2.5:1.5b",
        "prompt": f'Triage this customer support ticket:\n\n"{ticket}"',
        "format": TicketTriage.model_json_schema(),
        "stream": False,
    }, timeout=120)
    triage = TicketTriage.model_validate_json(resp.json()["response"])
    print(f"[{time.time() - start:.1f}s]", triage.model_dump())

Real, unedited output:

Ticket 1 [3.3s]: {'category': 'billing', 'severity': 'critical', 'sentiment': 'negative',
  'action_items': ['contact support team immediately to discuss the issue']}

Ticket 2 [3.0s]: {'category': 'account', 'severity': 'low', 'sentiment': 'neutral',
  'action_items': []}

Ticket 3 [3.4s]: {'category': 'technical', 'severity': 'critical', 'sentiment': 'negative',
  'action_items': ['check if there are any updates available for the app or software version',
  'contact technical support or developer of the application to investigate further']}

All three validated cleanly, with sensible, differentiated triage decisions: the double-charge billing complaint and the deadline-blocking crash both came back negative sentiment and critical severity, while the “not urgent” account question came back neutral and low severity with an empty action list. Note that action_items is allowed to be an empty list, that is valid per the schema (a plain list[str] has no minimum length), and it is the right answer for a ticket that genuinely needs no follow-up action yet.

Step 5: Know What the Schema Actually Enforces

Not every model is equally reliable at filling in that schema well, and not every JSON Schema keyword is actually enforced by Ollama. Both are worth checking directly instead of assuming.

Smaller models stay valid, but get sloppier

Run the exact same three tickets through qwen2.5:0.5b, a smaller model in the same family:

Ticket 1: {'category': 'billing', 'severity': 'critical', 'sentiment': 'neutral', 'action_items': []}
Ticket 2: {'category': 'account', 'severity': 'high', 'sentiment': 'neutral',
  'action_items': ['Contact your administrator or system admin to apply for the change of email
  address', "Verify that the new email address is unique and free from spam or phishing..."]}
Ticket 3: {'category': 'technical', 'severity': 'high', 'sentiment': 'neutral',
  'action_items': ['analyze problem, ']}

Every response here still validated successfully, the schema constraint held. But the content quality dropped noticeably: an angry double-charge complaint came back with sentiment: neutral, a “not urgent” question was rated high severity and given a wall of unnecessary security advice, and ticket 3’s action item is a truncated fragment with a stray trailing comma. Schema-valid and useful are not the same thing. If a task is complex enough to need real judgment, budget for a larger model, and treat a smaller model’s structurally-valid output with more skepticism, not less.

Array length and numeric range constraints are enforced

Pydantic’s Field() can add constraints beyond basic types, for example a minimum list length. Add Field(min_length=5) to action_items and Ollama actually honors it:

from pydantic import Field

class StrictTicket(BaseModel):
    category: str
    action_items: list[str] = Field(min_length=5)

Across four separate runs against a simple “change my email” ticket, every single response came back with 5 or more action items, satisfying the constraint every time. The same is true for numeric bounds: a rating: int = Field(ge=1, le=5) constraint reliably kept generated ratings inside that range in testing. Ollama compiles the schema into a generation grammar, and array length and numeric range are both part of that grammar.

Regex patterns are not enforced, they fail the request outright

Not every JSON Schema keyword survives that compilation step. Try adding a pattern constraint, a regular expression, to a field:

class TicketId(BaseModel):
    ticket_id: str = Field(pattern=r"^TICKET-\d{4}$")
    category: str

resp = requests.post(OLLAMA_URL, json={
    "model": MODEL,
    "prompt": "Make up a ticket ID and category for a billing complaint.",
    "format": TicketId.model_json_schema(),
    "stream": False,
})
print(resp.status_code, resp.text)

Real output:

400
{"error":"{\"error\":{\"code\":400,\"message\":\"Failed to initialize samplers: failed to parse
grammar\",\"type\":\"invalid_request_error\"}}"}

This is not a silent failure, and that is actually the better outcome: the whole request is rejected with an HTTP 400 before any generation happens, so you find out immediately in development rather than shipping a schema that quietly does nothing. The practical takeaway is to keep your Pydantic field constraints to the kinds JSON Schema-to-grammar compilation reliably supports (types, enums, required fields, array length, numeric ranges) and enforce anything regex-shaped, like a specific ID format, with your own Pydantic field_validator after the fact instead of leaning on pattern.

Step 6: The Reasoning Model Gotcha

So far every example used qwen2.5:1.5b, which is not a reasoning model. Run the identical ticket-triage schema against qwen3.5:4b, which is, and something breaks:

Ticket 1 [19.6s]: raw response: ''
VALIDATION FAILED:
1 validation error for TicketTriage
  Invalid JSON: EOF while parsing a value at line 1 column 0

An empty string, every time, for every ticket. To see why, request the same call with "stream": False and print the full response object instead of just the response field:

{
  "response": "",
  "thinking": "{\n    \"category\": \"[1]\",\n    \"severity\": \"[2]\"\n}",
  "done": true,
  "done_reason": "stop",
  ...
}

There it is. qwen3.5:4b is a reasoning model: by default it generates an internal “thinking” pass before its real answer, exposed separately in the thinking field. Here, the schema-shaped content ended up inside that internal thinking field instead of the actual response, and the model stopped before producing anything in response itself. Your application only ever sees the empty response field, so this looks exactly like the model silently failing.

The fix

Ollama’s /api/generate endpoint accepts a think parameter for reasoning models. Set it to false to skip the internal thinking pass:

resp = requests.post(OLLAMA_URL, json={
    "model": "qwen3.5:4b",
    "prompt": prompt,
    "format": TicketTriage.model_json_schema(),
    "think": False,
    "stream": False,
})

Real output with the fix applied, run against the same three tickets from Step 4:

Ticket 1 [15.3s]: {'category': 'billing', 'severity': 'high', 'sentiment': 'negative',
  'action_items': ['Immediately verify transaction records to locate the duplicate charge.',
  'Contact payment method directly or issue an instant refund/cancellation for the erroneous
  amount.', "Acknowledge customer's frustration and apologize for lack of communication on
  previous emails."]}

Ticket 2 [6.8s]: {'category': 'technical', 'severity': 'low', 'sentiment': 'neutral',
  'action_items': ['Answer how to update the email address in account settings.',
  'Follow up if additional support is needed.']}

Ticket 3 [8.9s]: {'category': 'technical', 'severity': 'critical', 'sentiment': 'negative',
  'action_items': ['Identify specific report types triggering the crash.', 'Analyze server logs
  for resource bottlenecks or data corruption issues during export generation.', 'Provide a
  temporary workaround (e.g., batch processing, CSV instead of interactive reports) to unblock
  user immediately.']}

All three valid, and notably more thorough than the smaller model’s output from Step 5. (Ticket 2 landed on technical instead of account this time, a reminder that category choice on an ambiguous ticket is a judgment call even for a capable model, not something the schema itself controls.) If you are using a reasoning model with structured outputs and getting mysterious empty responses, this is very likely why. Check the thinking field in the raw response before assuming anything else is wrong.

Step 7: Wrap It in a Retry Helper for Production

Everything above ran cleanly on the first try, but real network calls fail sometimes, and it is worth having a real retry path rather than discovering you need one during an incident. Here is a small helper that retries on both request errors and validation failures, and defaults think to False so it does not fall into Step 6’s trap by default:

import time
from typing import Type, TypeVar
import requests
from pydantic import BaseModel, ValidationError

OLLAMA_URL = "http://localhost:11434/api/generate"
T = TypeVar("T", bound=BaseModel)


class OllamaExtractionError(Exception):
    pass


def extract_structured(
    model: str,
    prompt: str,
    schema_model: Type[T],
    max_retries: int = 3,
    think: bool = False,
) -> T:
    last_error = None
    for attempt in range(1, max_retries + 1):
        try:
            resp = requests.post(
                OLLAMA_URL,
                json={
                    "model": model,
                    "prompt": prompt,
                    "format": schema_model.model_json_schema(),
                    "think": think,
                    "stream": False,
                },
                timeout=120,
            )
            resp.raise_for_status()
            raw_text = resp.json()["response"]
            return schema_model.model_validate_json(raw_text)
        except (requests.exceptions.RequestException, ValidationError) as e:
            last_error = e
            print(f"  attempt {attempt}/{max_retries} failed: {e.__class__.__name__}: {e}")
            time.sleep(1)
    raise OllamaExtractionError(
        f"Failed to get valid {schema_model.__name__} after {max_retries} attempts"
    ) from last_error

Watch it succeed

result = extract_structured(
    model="qwen2.5:1.5b",
    prompt="Triage this customer support ticket:\n\n\"I've been charged twice...\"",
    schema_model=TicketTriage,
)
print(result.model_dump())
# {'category': 'billing', 'severity': 'critical', 'sentiment': 'negative',
#  'action_items': ['contact support team immediately to discuss the issue']}

Watch it fail honestly, and see why retries are not a fix for everything

Call the same helper against qwen3.5:4b with think=True, reproducing Step 6’s bug on purpose:

extract_structured(
    model="qwen3.5:4b",
    prompt="Triage this ticket: I want a refund now.",
    schema_model=TicketTriage,
    max_retries=3,
    think=True,
)
  attempt 1/3 failed: ValidationError: 1 validation error for TicketTriage
  Invalid JSON: EOF while parsing a value at line 1 column 0
  attempt 2/3 failed: ValidationError: 1 validation error for TicketTriage
  Invalid JSON: EOF while parsing a value at line 1 column 0
  attempt 3/3 failed: ValidationError: 1 validation error for TicketTriage
  Invalid JSON: EOF while parsing a value at line 1 column 0
OllamaExtractionError: Failed to get valid TicketTriage after 3 attempts

All three attempts failed identically. This is the honest, useful behavior of a retry loop: it protects you against transient issues (a dropped connection, a momentary hiccup), but it cannot fix a systematic misconfiguration, the same request will just fail the same way every time. Three identical failures in a row is a signal to go check your setup, in this exact case, to check whether think is set correctly, not a signal to raise max_retries and hope.

Common Mistakes and How to Verify Your Setup

  • Trusting format: "json" alone as a contract. It guarantees syntactically valid JSON, not any particular set of field names. Verify by running the same prompt three or four times and diffing the key names; if they are not identical every time, you need a real schema, not just JSON mode.
  • Assuming schema-valid means factually correct. Every example in Step 3 called a two-week shipping delay “no complaint,” even under a strictly enforced schema. Validate structure with Pydantic; validate substance by spot-checking real outputs against the source text, especially before trusting a field like a sentiment or severity label downstream.
  • Getting silent empty responses from a reasoning model. If response is empty but the call succeeded (HTTP 200, done: true), check the thinking field in the raw JSON before assuming Ollama or your schema is broken. Set "think": false if you do not need the reasoning trace.
  • Using Pydantic constraints that JSON Schema-to-grammar compilation cannot express. pattern (regex) causes an outright HTTP 400 in this version of Ollama, “failed to parse grammar.” Verify any constraint you rely on by actually sending it and checking the response code, do not assume every Field() option is enforced just because model_json_schema() included it.
  • Trusting a small model with a complex schema. Structural validity does not imply good judgment, as Step 5’s 0.5B-parameter comparison shows directly. If a schema has more than two or three fields, or any field that requires real reasoning about the input, test with your actual target model before assuming a smaller, faster one will do.

To confirm your own setup end to end, run the retry helper from Step 7 against a real prompt and a model you plan to use, and check three things in the result: that it returns a real Pydantic instance (not a string you still need to parse), that the field values are substantively correct for your input (not just structurally present), and that a deliberately broken input (an empty prompt, or a schema field that cannot reasonably be filled) fails loudly through OllamaExtractionError rather than silently returning nonsense.

Next Steps

From here, a few natural directions to take this:

  • If you are building a full agent rather than a single extraction call, Pydantic AI gives you the same guaranteed-valid-instance outcome as extract_structured() above, plus tool registration, dependency injection, and multi-turn state, without you having to hand-write the retry and validation plumbing yourself. See How to Build a Production-Grade AI Agent With Pydantic AI and Ollama for a full walkthrough.
  • If your extraction schema needs to route between models by task difficulty, for example a cheap model for simple triage and a larger one when the schema or input is more complex, see How to Build a Multi-Tier LLM Router With Automatic Fallback Using Python and Ollama.
  • Try adding Field(description="...") to individual Pydantic fields. Descriptions are included in the generated JSON Schema and give the model extra context about what each field means, which can measurably improve accuracy on ambiguous fields like severity or sentiment without changing your code’s structure at all.
  • For schemas with nested objects or lists of objects rather than lists of strings, the same mechanism applies: nest one Pydantic model inside another, and model_json_schema() handles the $defs and references automatically. Test nested schemas the same way this tutorial tested flat ones, by running real prompts several times and checking both validity and content quality before trusting them.

Tags:

Data ValidationJSON SchemaLocal AIOllamaPython

Share

A municipal water tower in Minnesota against an overcast sky, representing the type of water infrastructure targeted in the attacks
Previous Post

Iran’s Water Utility Hacks Turn Exposed PLCs Into a National Security Problem

A red emergency life ring safety station mounted on a post beside a marina dock
Next Post

OpenAI Disbands Its Preparedness Team Amid a Wave of Safety Departures

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
08 Oct
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
08 Oct
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
Trending
October 8, 2026
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
October 8, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
October 8, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google

Related Posts

A laptop wrapped in a chain and padlock, illustrating least-privilege controls for AI agents.
Learning Hub

How to Secure Tool-Using AI Agents Before They Touch Production

June 8, 2026
Colorful sticky notes arranged on an office wall, symbolizing governance checklists and planning.
Learning Hub

AI Governance for Agentic Apps: A Practical Checklist for Builders

June 8, 2026
A technician connects green fiber optic cables at a data center, representing a private production inference endpoint.
Learning Hub

How to Deploy a Fine-Tuned LLM Behind a Private Production Inference Endpoint

June 8, 2026
Narrow aisle behind black supercomputer racks in a data center
Learning Hub

Kubernetes SELinux Volume Labeling: What Cluster Operators Should Audit Before v1.37

June 8, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026