How to Build a Multi-Tier LLM Router With Automatic Fallback Using Python and Ollama
A step-by-step tutorial for routing prompts to different local models by complexity, with automatic fallback when the primary model errors out or times out, built and tested end to end with Ollama.
Most tutorials that call a language model send every request to the same model. That works until the bill arrives or the provider has a bad day. A recent freeCodeCamp tutorial lays out a common fix: classify each incoming prompt by how hard it looks, route simple requests to a small, cheap model and hard ones to a bigger one, and automatically retry on a different model when the first choice fails. That third part, automatic fallback, is where most first attempts quietly break, because the code only watches for one kind of failure and misses the others.
Table Of Content
- What tiered routing and fallback actually mean
- Prerequisites
- Step 1: Install Ollama and pull three differently sized models
- Step 2: Set up the Python project
- Step 3: Classify prompt complexity with a deterministic analyzer
- Step 4: Route each tier to a real model over Ollama’s HTTP API
- Why the complex tier times out
- Fix it with a per-tier timeout
- Step 5: Handle a parameter only some models support
- Why this breaks the two smaller tiers
- Fix it by tracking which models support thinking
- Step 6: Add automatic fallback for real failures
- The mistake hiding in that except clause
- Step 7: Deliberately test the fallback path
- Step 8: Run the complete pipeline
- Common mistakes and gotchas
- How to verify everything works end to end
- Next steps
This tutorial builds that pattern yourself, end to end, in Python, using Ollama to run three real local models of different sizes for free, with no API key and no cloud bill. Every piece of code below, including two real bugs, was written and run against a live local Ollama installation while preparing this tutorial, not copied from a working example. You will hit both bugs yourself, see the exact error each one produces, and fix them the same way they were fixed here.
What tiered routing and fallback actually mean
A large language model (LLM) is a program that generates text in response to a prompt. Different LLMs trade off cost, speed, and capability: a small model answers in a second and costs almost nothing to run, but struggles with genuinely hard requests; a large model handles those hard requests well but is slower and, on a paid API, far more expensive per request.
Tiered routing means looking at each incoming prompt, deciding how difficult it is, and sending it to the smallest model that can handle it, rather than sending every request to the same model regardless of difficulty. Fallback is a separate, second safety net: if the model a request gets routed to fails for any reason (it does not exist, the connection drops, it takes too long to answer), the request is automatically retried against a different, working model, instead of the caller just getting an exception.
These two ideas solve different problems and both matter. Tiered routing controls cost and latency in the normal case. Fallback controls what happens when the normal case breaks. A router with tiering but no fallback still goes down the moment one model in the chain has a problem.
Prerequisites
- Operating system: commands below are written for Ubuntu 24.04 LTS or newer (Debian-based Linux in general). Ollama also ships macOS and Windows builds; only the install command in Step 1 changes on those platforms.
- Hardware: at least 8 GB of RAM and roughly 5 GB of free disk space for the three models this tutorial pulls (about 400 MB, 1 GB, and 3.4 GB). No GPU is required; all three models run fine on CPU, though the largest one is noticeably slower.
- Python: Python 3.9 or newer, with
python3 -m venvavailable (it ships with the standard library on any normal Python install). - Internet access: required once, to install Ollama and download the three models. Everything after that runs fully offline.
- Prior knowledge assumed: comfort running commands in a terminal and reading basic Python, including
try/except. No prior experience with Ollama or LLM routing is assumed.
Step 1: Install Ollama and pull three differently sized models
Ollama is a tool for downloading and running open-weight language models directly on your own machine, so every prompt you send in this tutorial stays local. Install it with the official script:
# Context: Ubuntu 24.04 LTS or newer, sudo access.
# Purpose: download and install Ollama, and set it up as a background service.
curl -fsSL https://ollama.com/install.sh | sh
Expected output: the script detects your OS, downloads the Ollama binary, and configures it as a systemd service that starts automatically. Confirm it installed correctly:
ollama --version
Expected output: a line similar to ollama version is 0.31.2. This tutorial was written and tested against that exact version.
Now pull three models that stand in for three tiers of capability. Using the same model family for the two smaller tiers, Alibaba’s Qwen 2.5, keeps the comparison clean; the largest tier uses the newer Qwen 3.5, which Ollama’s own model library tags as a “thinking” model, meaning it can reason step by step before answering. That distinction matters later.
# Purpose: download the small, fast model for simple requests (about 400 MB).
ollama pull qwen2.5:0.5b
# Purpose: download the mid-sized model for moderate requests (about 1 GB).
ollama pull qwen2.5:1.5b
# Purpose: download the reasoning-capable model for hard requests (about 3.4 GB).
ollama pull qwen3.5:4b
Expected output: three separate progress bars, each ending with success. Confirm all three are available locally:
ollama list
Expected output: a table listing qwen2.5:0.5b, qwen2.5:1.5b, and qwen3.5:4b with their size and pull date. If a tag is missing here, any router code that references it will fail with a 404 “model not found” error, which you will deliberately trigger yourself in Step 7, so it is worth confirming all three rows are present now.
Step 2: Set up the Python project
# Purpose: create and activate an isolated Python environment for this tutorial.
mkdir llm-router-tutorial && cd llm-router-tutorial
python3 -m venv venv
source venv/bin/activate
Expected output: your shell prompt gains a (venv) prefix. Every pip and python3 command from here on should run inside this same activated terminal. Install the one dependency this tutorial needs:
pip install requests
This tutorial calls Ollama’s native HTTP API directly with the requests library, rather than a wrapper library, so every request and every failure mode is visible in plain Python instead of hidden behind an abstraction.
Step 3: Classify prompt complexity with a deterministic analyzer
Before a request can be routed anywhere, something has to decide how hard it is. Calling an LLM just to make that decision would add cost and latency to every single request, so this step uses a fast, deterministic (rule-based, not model-based) classifier instead: it looks at word count, whether the prompt contains a code block, and whether it contains any of a short list of keywords that tend to signal a harder task. Create prompt_analyzer.py:
import re
from enum import Enum
class TaskComplexity(Enum):
SIMPLE = "simple"
MEDIUM = "medium"
COMPLEX = "complex"
class PromptAnalyzer:
def __init__(self):
self.complex_keywords = [
r"\brefactor\b",
r"\bdebug\b",
r"\bwrite code\b",
r"\banalyze\b",
r"\balgorithm\b",
r"\barchitecture\b",
]
def analyze_complexity(self, prompt: str) -> TaskComplexity:
normalized = prompt.lower().strip()
word_count = len(normalized.split())
contains_code = "```" in prompt
has_complex_keyword = any(
re.search(pattern, normalized) for pattern in self.complex_keywords
)
if contains_code or has_complex_keyword or word_count > 40:
return TaskComplexity.COMPLEX
elif word_count > 12:
return TaskComplexity.MEDIUM
else:
return TaskComplexity.SIMPLE
if __name__ == "__main__":
analyzer = PromptAnalyzer()
tests = [
"What time zone is UTC?",
"Summarize the difference between TCP and UDP in a couple of sentences for a junior developer.",
"Refactor this function to use a dictionary lookup instead of a chain of if statements and explain the algorithmic complexity change.",
]
for t in tests:
print(f"{analyzer.analyze_complexity(t).value:8} | {t}")
Run it:
python3 prompt_analyzer.py
Expected output (this is real, captured output, not a mockup):
simple | What time zone is UTC?
medium | Summarize the difference between TCP and UDP in a couple of sentences for a junior developer.
complex | Refactor this function to use a dictionary lookup instead of a chain of if statements and explain the algorithmic complexity change.
The short question gets classified simple because it is under 12 words. The TCP/UDP question lands on medium because it is longer than 12 words but does not hit any complex trigger. The refactor request hits complex purely on the keyword check: at 21 words it does not come close to the 40-word cutoff on its own, but it matches the \brefactor\b pattern, and that alone is enough to mark it complex regardless of length. These thresholds (12 and 40 words) are tuned for this tutorial’s short example prompts; in a real application, set them by looking at your own traffic’s actual word-count distribution, not by guessing.
Step 4: Route each tier to a real model over Ollama’s HTTP API
With classification working, the next step maps each tier to an actual model and calls it. Create model_router.py with a first version:
import requests
from prompt_analyzer import PromptAnalyzer, TaskComplexity
OLLAMA_CHAT_URL = "http://localhost:11434/api/chat"
TIER_MODELS = {
TaskComplexity.SIMPLE: "qwen2.5:0.5b",
TaskComplexity.MEDIUM: "qwen2.5:1.5b",
TaskComplexity.COMPLEX: "qwen3.5:4b",
}
class ModelRouter:
def __init__(self, chat_url=OLLAMA_CHAT_URL, timeout=60):
self.chat_url = chat_url
self.timeout = timeout
self.analyzer = PromptAnalyzer()
def _call_model(self, model: str, prompt: str) -> str:
response = requests.post(
self.chat_url,
json={
"model": model,
"messages": [{"role": "user", "content": prompt}],
"stream": False,
},
timeout=self.timeout,
)
response.raise_for_status()
return response.json()["message"]["content"]
def route(self, prompt: str) -> dict:
tier = self.analyzer.analyze_complexity(prompt)
model = TIER_MODELS[tier]
content = self._call_model(model, prompt)
return {"tier": tier.value, "model_used": model, "content": content}
The /api/chat endpoint is Ollama’s native REST API for conversational requests: it takes a model name and a list of chat-style messages, and with "stream": False it waits for the full response instead of streaming it back token by token. A 60-second timeout on the HTTP request is a reasonable-looking default, so test it with all three tiers:
python3 -c "
from model_router import ModelRouter
router = ModelRouter()
print(router.route('Refactor this function to use a dictionary lookup instead of a chain of if statements and explain the algorithmic complexity change.'))
"
Expected output: this is the first real bug. On this test machine, the call to qwen3.5:4b raised this exact exception after waiting the full 60 seconds:
requests.exceptions.ReadTimeout: HTTPConnectionPool(host='localhost', port=11434): Read timed out. (read timeout=60)
Why the complex tier times out
Ollama’s own model library tags qwen3.5 as a thinking model: before writing its final answer, it works through an internal reasoning process, which takes real, measurable time and produces a noticeably longer response. A 500 million or 1.5 billion parameter non-reasoning model answers a short prompt in a few seconds; a 4 billion parameter reasoning model working through the same prompt can easily take over a minute. A single fixed timeout tuned for the fast tiers will routinely cut off the slow, capable tier before it finishes, which is exactly backward: the hardest requests are the ones a timeout should least want to abandon.
Fix it with a per-tier timeout
Update model_router.py so each tier carries its own timeout budget instead of sharing one:
TIER_TIMEOUTS = {
TaskComplexity.SIMPLE: 30,
TaskComplexity.MEDIUM: 60,
TaskComplexity.COMPLEX: 180,
}
And change route to look up the timeout per tier instead of using a fixed value:
def route(self, prompt: str) -> dict:
tier = self.analyzer.analyze_complexity(prompt)
model = TIER_MODELS[tier]
timeout = TIER_TIMEOUTS[tier]
content = self._call_model(model, prompt, timeout)
return {"tier": tier.value, "model_used": model, "content": content}
(and add the matching timeout parameter to _call_model‘s signature, passed through to requests.post). Re-run the same test. On this machine, repeated runs of the identical prompt against qwen3.5:4b finished anywhere from about a minute and a half to two and a half minutes, comfortably inside the new 180-second budget every time this was tested in isolation. Those numbers are not a universal constant, they depend entirely on your CPU, and the run-to-run variance itself is worth noticing now: it is large enough that a busier machine, or simply an unlucky run, can push the same request past even a generous timeout. Step 8 catches exactly that happening. The lesson is the shape of the fix, not the specific 180: size your timeout budget per tier based on what you actually observe that tier taking, with headroom, not a single guess applied everywhere.
Step 5: Handle a parameter only some models support
Ollama’s API documentation describes an optional think field on /api/chat: “(for thinking models) should the model think before responding?” That parenthetical matters more than it looks. Suppose you want to explicitly confirm thinking is enabled for the complex tier, so you add it to every request:
response = requests.post(
self.chat_url,
json={
"model": model,
"messages": [{"role": "user", "content": prompt}],
"stream": False,
"think": True,
},
timeout=timeout,
)
Why this breaks the two smaller tiers
Testing this change against all three tiers reproduces a second real bug: both qwen2.5:0.5b and qwen2.5:1.5b immediately return requests.exceptions.HTTPError: 400 Client Error: Bad Request for url: http://localhost:11434/api/chat. The response body is exactly as direct as the exception: {"error":"\"qwen2.5:1.5b\" does not support thinking"}. That is not a guess about what went wrong, it is Ollama itself confirming it: think only works on models built for it, and qwen2.5 is not one of them, matching Ollama’s library page, which marks qwen3.5 with a thinking capability tag while qwen2.5 carries no such tag. On this Ollama version the request is rejected outright rather than silently ignored. A router that assumes every model accepts the same request shape will break the instant it adds a tier running a different model family.
Fix it by tracking which models support thinking
Track reasoning capability per model and only include think when it is actually supported:
REASONING_MODELS = {"qwen3.5:4b"}
def _call_model(self, model: str, prompt: str, timeout: int) -> str:
payload = {
"model": model,
"messages": [{"role": "user", "content": prompt}],
"stream": False,
}
if model in REASONING_MODELS:
payload["think"] = True
response = requests.post(self.chat_url, json=payload, timeout=timeout)
response.raise_for_status()
return response.json()["message"]["content"]
Re-running all three tiers after this change succeeds without a 400 anywhere. The general principle: in any router that fans out to more than one model, treat “what parameters does this specific model accept” as part of that model’s configuration, never assume it is universal.
Step 6: Add automatic fallback for real failures
Per-tier timeouts and per-model parameters fix two specific bugs, but production traffic eventually hits failures neither one prevents: a model tag gets removed, the connection drops mid-request, or a provider has an outage. This is where fallback belongs, wrap the primary call in a try/except and retry on a known-good model if it fails:
FALLBACK_MODEL = "qwen2.5:1.5b"
def route(self, prompt: str) -> dict:
tier = self.analyzer.analyze_complexity(prompt)
primary_model = TIER_MODELS[tier]
timeout = TIER_TIMEOUTS[tier]
try:
content = self._call_model(primary_model, prompt, timeout)
return {
"tier": tier.value,
"model_requested": primary_model,
"model_used": primary_model,
"used_fallback": False,
"content": content,
}
except requests.exceptions.RequestException as exc:
fallback_content = self._call_model(FALLBACK_MODEL, prompt, timeout=60)
return {
"tier": tier.value,
"model_requested": primary_model,
"model_used": FALLBACK_MODEL,
"used_fallback": True,
"error": f"{type(exc).__name__}: {exc}",
"content": fallback_content,
}
The mistake hiding in that except clause
It is tempting to write except requests.exceptions.HTTPError, since that is what catches an HTTP 404 or 400. Step 4’s timeout, however, raised requests.exceptions.ReadTimeout, and a dropped connection raises requests.exceptions.ConnectionError. Neither one is a subclass of HTTPError in the requests library; both are siblings of it under the shared base class requests.exceptions.RequestException. A router that only catches HTTPError will correctly fall back on a 404 “model not found” but will let a hung or slow request crash the whole call with an unhandled exception instead of failing over, which defeats the point of building fallback in the first place. Catching RequestException, the shared base class, covers both failure families in one except.
Step 7: Deliberately test the fallback path
Trusting fallback logic without seeing it actually fire is a bad habit. Simulate a provider outage by pointing a tier at a model tag that does not exist:
from model_router import ModelRouter, TIER_MODELS
from prompt_analyzer import TaskComplexity
router = ModelRouter()
TIER_MODELS[TaskComplexity.SIMPLE] = "qwen2.5:0.5b-does-not-exist"
result = router.route("What time zone is UTC?")
print(f"model_requested={result['model_requested']}")
print(f"model_used={result['model_used']}")
print(f"used_fallback={result['used_fallback']}")
print(f"error={result.get('error')}")
Expected output (real, captured output):
model_requested=qwen2.5:0.5b-does-not-exist
model_used=qwen2.5:1.5b
used_fallback=True
error=HTTPError: 404 Client Error: Not Found for url: http://localhost:11434/api/chat
Ollama returned a genuine 404 for the nonexistent tag, the except block caught it, and the router transparently retried on qwen2.5:1.5b and returned a real answer. This is the entire point of the pattern: the caller gets a working response and a clear record of what happened, instead of a stack trace.
Step 8: Run the complete pipeline
Create app.py to exercise all three tiers together:
from model_router import ModelRouter
TEST_PROMPTS = [
"What time zone is UTC?",
"Summarize the difference between TCP and UDP in a couple of sentences for a junior developer.",
"Refactor this function to use a dictionary lookup instead of a chain of if statements and explain the algorithmic complexity change.",
]
def main():
router = ModelRouter()
for prompt in TEST_PROMPTS:
result = router.route(prompt)
print(f"prompt: {prompt}")
print(
f" tier={result['tier']} "
f"model_used={result['model_used']} "
f"used_fallback={result['used_fallback']}"
)
print(f" response: {result['content'][:200].strip()}...")
print()
if __name__ == "__main__":
main()
python3 app.py
Expected output (real, captured output from this exact run):
prompt: What time zone is UTC?
tier=simple model_used=qwen2.5:0.5b used_fallback=False
response: UTC stands for Coordinated Universal Time...
prompt: Summarize the difference between TCP and UDP in a couple of sentences for a junior developer.
tier=medium model_used=qwen2.5:1.5b used_fallback=False
response: TCP (Transmission Control Protocol) ensures reliable delivery of data...
prompt: Refactor this function to use a dictionary lookup instead of a chain of if statements and explain the algorithmic complexity change.
tier=complex model_used=qwen2.5:1.5b used_fallback=True
Notice the third line: on this particular run, the complex tier itself triggered a fallback, even with the 180-second budget from Step 4. Running that identical prompt against qwen3.5:4b alone, moments later, succeeded in about two and a half minutes with no fallback at all. Nothing about the code changed between those two runs; the reasoning model’s own response time varied enough, run to run, to cross a fixed 180-second line on one attempt and stay under it on another. That variance is not a bug to chase down, it is the real, honest behavior of a reasoning model under CPU inference, and it is precisely why fallback earns its place even after timeouts are tuned carefully: you cannot fully bound a reasoning model’s worst-case latency in advance, so the system needs a graceful path for when an individual request runs long, not just a bigger number.
Common mistakes and gotchas
- Catching only
HTTPError: misses timeouts and connection errors entirely. Catchrequests.exceptions.RequestException, the shared base class, unless you deliberately need different handling per failure type. - One timeout for every tier: a value that comfortably fits a small model will cut off a slower, more capable one before it finishes. Set a timeout per tier based on what that tier actually takes, with headroom, not a single guess.
- Assuming every model accepts the same request fields: parameters like
thinkare model-family-specific. Track capability per model rather than adding a field to every request. - Word-count thresholds copied verbatim: the 12- and 40-word cutoffs here fit this tutorial’s short test prompts. Calibrate against your own real traffic’s word-count distribution before trusting a classifier in production.
- An unbounded fallback chain: this tutorial falls back exactly one hop, to a single known-good model. In production, cap the total number of attempts and log every fallback event; a fallback that fires on every request is masking a chronic problem with the primary model, not fixing one.
How to verify everything works end to end
- Run
ollama listand confirm all three models,qwen2.5:0.5b,qwen2.5:1.5b, andqwen3.5:4b, are present. - Run
python3 app.pyand confirm three prompts print three tier labels (simple, medium, complex) with real response text under each, not a traceback. - Deliberately break a tier the way Step 7 does, by pointing
TIER_MODELSat a nonexistent tag, and confirmused_fallbackcomes backTruewith a real 404 captured inerror, and that the response text is still a genuine answer. - Restore the correct tag afterward and re-run once more to confirm the primary path still works with
used_fallback=False.
Next steps
The router built here is provider-agnostic on purpose: swapping Ollama for a paid API means changing only the code inside _call_model, since the tiering, timeout, and fallback logic around it does not depend on which provider serves the request. From here, consider logging the tier, model used, and whether fallback fired for every request, so you can see cost and reliability trends over time instead of guessing at them. A circuit breaker that temporarily stops sending requests to a primary model after repeated failures, rather than retrying it on every single call during an outage, is a natural next layer once basic fallback is working. And if the reasoning tier’s cold-start or variable latency becomes a bottleneck under real load, look at keeping it warm or batching requests to it, rather than only widening the timeout further.








No Comment! Be the first one.