How to Cut AI Agent Latency With Streaming Responses and Parallel Tool Calls in Python
Learn to measure and fix two different causes of slow-feeling AI agents: streaming responses for faster perceived replies, and parallel tool calls to cut real wall-clock wait time.
Two AI agents can run the exact same model and still feel completely different to use. One seems to think forever before saying a word. The other starts talking almost immediately, even when it takes just as long to finish the full answer. The difference usually has nothing to do with the model itself: it comes from how the agent’s code handles two separate things, how quickly the first piece of an answer reaches the user, and how much time the agent wastes waiting on tool calls it could have run at the same time instead of one after another.
Table Of Content
- What Is Time to First Token, and Why Does It Need a Different Fix Than Total Time?
- What You Will Accomplish
- Prerequisites
- Step 1: Install Ollama and Pull a Model
- A Gotcha Worth Knowing Before You Start Measuring
- Step 2: Establish a Non-Streaming Baseline
- Step 3: Switch to Streaming and Watch TTFT Drop
- Step 4: Prove Longer Context Costs More at the Front, Not the Back
- Step 5: Measure the Cost of Sequential Tool Calls
- Step 6: Parallelize Independent Tool Calls
- Step 7: Put It Together: A Streaming Agent With Parallel Tool Calls
- Common Mistakes and Gotchas
- How to Verify It All Works End to End
- Next Steps
This tutorial teaches you to measure both problems and fix them. You will learn the difference between time to first token (TTFT) and total generation time, why they need different fixes, and how running independent tool calls at the same time instead of sequentially can cut an agent’s wall-clock time dramatically. Every number in this tutorial comes from commands you run yourself against a real local model, not from a spec sheet.
What Is Time to First Token, and Why Does It Need a Different Fix Than Total Time?
When you send a prompt to a language model, two things happen in sequence. First, the model reads and processes your entire input, the system prompt, the conversation history, and the new message, before it can produce a single word of output. This reading step is called prefill. Only after prefill finishes can the model start generating output tokens one at a time.
Time to first token (TTFT) is how long a user waits before seeing anything at all. It is dominated by prefill, plus a small amount of overhead. Total generation time is how long the entire response takes to finish, prefill plus the time to generate every output token.
These two numbers respond to different fixes. Shrinking the size of your prompt and conversation history reduces TTFT, because there is less to read before the model can start. Asking for a shorter, more structured answer reduces total generation time, because there are fewer tokens to produce. Streaming the response instead of waiting for the whole thing does not change either number on the server. What it changes is what your client can show the user: instead of blocking until everything is ready, you display each token as it arrives, so the user perceives the TTFT delay instead of the total delay.
What You Will Accomplish
- Measure real TTFT and total generation time for both a blocking (non-streaming) response and a streamed response from a local model
- See exactly why streaming makes an agent feel dramatically faster, even when the total time barely changes
- Prove empirically that a longer prompt increases TTFT (prefill cost), independent of how long the answer itself is
- Build a small tool-calling agent, measure how much time sequential tool calls waste, then fix it with parallel execution
- Combine both techniques, streaming and parallel tool calls, into one working agent and verify the timing adds up the way you would expect
Prerequisites
- Windows, macOS, or Linux with about 4 GB of free disk space for a small local model
- Ollama, a tool for running open-weight language models on your own machine. This tutorial installs and tests everything against a model running locally, so no cloud account or API key is needed.
- Python 3.10 or later, with the
requestspackage installed (pip install requests) - Basic comfort running commands in a terminal and reading Python. No prior experience with threading or async code is required; it is explained as it comes up.
Step 1: Install Ollama and Pull a Model
Download Ollama from ollama.com/download for your platform and run the installer. On Windows this is a normal graphical installer; on macOS and Linux, follow the platform-specific instructions on that page. Once it is installed, Ollama runs as a background service that listens on http://localhost:11434.
Pull a small, tool-calling-capable model:
ollama pull qwen3.5:4b
This downloads a 3.7 GB model. Expect output like this while it downloads:
pulling manifest
pulling 81fb60c7daa8: 100% |███████████████| 3.4 GB/3.4 GB
verifying sha256 digest
writing manifest
success
Confirm it is available:
ollama list
NAME ID SIZE MODIFIED
qwen3.5:4b 2a654d98e6fb 3.7 GB ...
qwen3.5:4b is used throughout this tutorial because it is small enough to run on modest hardware and supports the tool-calling format used in the later steps. Any similarly-sized Ollama model that supports tools will work the same way; the exact timing numbers you see will depend on your own CPU and GPU.
A Gotcha Worth Knowing Before You Start Measuring
The first version of the baseline script below was run without any special settings, on the same three-sentence prompt used in Step 2. The visible answer was three sentences, about 57 words, but the response reported eval_count: 1185, generating over a thousand tokens for a three-sentence answer. The reason is that qwen3.5 is a reasoning model: by default it generates an internal chain of thought before writing its visible answer, and that hidden reasoning counts toward eval_count and eval_duration even though you never see it in message.content. That first run took nearly two minutes for a three-sentence answer.
Every timed command in this tutorial passes "think": false in the request body, which turns that hidden reasoning off, so the numbers you see reflect the visible answer only. If your own numbers come out much larger than the ones shown here, check whether thinking is enabled; a good tell is an eval_count far larger than the word count of the answer you can see.
Step 2: Establish a Non-Streaming Baseline
Create 01_baseline.py:
import time
import requests
OLLAMA_URL = "http://localhost:11434/api/chat"
MODEL = "qwen3.5:4b"
PROMPT = "In exactly three sentences, explain what a race condition is."
payload = {
"model": MODEL,
"messages": [{"role": "user", "content": PROMPT}],
"stream": False,
"think": False,
}
start = time.perf_counter()
resp = requests.post(OLLAMA_URL, json=payload, timeout=120)
wall_time = time.perf_counter() - start
data = resp.json()
ns_to_s = 1_000_000_000
print(f"Client-observed wall time: {wall_time:.3f}s")
print(f"Server prompt_eval_duration (prefill): {data['prompt_eval_duration']/ns_to_s:.3f}s")
print(f"Server eval_duration (generation): {data['eval_duration']/ns_to_s:.3f}s")
print(f"Tokens generated: {data['eval_count']}")
Run it:
python 01_baseline.py
Real output from this tutorial’s test system (a workstation GPU shared with the CPU for this model; your numbers will vary with your own hardware):
Client-observed wall time: 7.428s
Server prompt_eval_duration (prefill): 0.608s
Server eval_duration (generation): 4.335s
Tokens generated: 75
With stream=False, requests.post() blocks until the entire response is ready. There is no way for your code to show the user anything before that 7.428 second mark: the client-observed wall time and the time to first visible content are the same number. Notice also that prompt_eval_duration (0.608s) plus eval_duration (4.335s) do not quite add up to the full wall time; the remainder is model load time and HTTP overhead, both of which the streaming version in the next step will also pay.
Step 3: Switch to Streaming and Watch TTFT Drop
Ollama’s streaming responses are newline-delimited JSON: each line is a complete JSON object representing one chunk, and the last line has "done": true along with the same timing fields you saw above. Create 02_streaming.py:
import json
import time
import requests
OLLAMA_URL = "http://localhost:11434/api/chat"
MODEL = "qwen3.5:4b"
PROMPT = "In exactly three sentences, explain what a race condition is."
payload = {
"model": MODEL,
"messages": [{"role": "user", "content": PROMPT}],
"stream": True,
"think": False,
}
start = time.perf_counter()
first_chunk_time = None
with requests.post(OLLAMA_URL, json=payload, stream=True, timeout=120) as resp:
for line in resp.iter_lines():
if not line:
continue
chunk = json.loads(line)
content = chunk.get("message", {}).get("content", "")
if content and first_chunk_time is None:
first_chunk_time = time.perf_counter()
total_time = time.perf_counter() - start
print(f"TTFT (first visible token): {first_chunk_time - start:.3f}s")
print(f"Total time (last token): {total_time:.3f}s")
Run it:
python 02_streaming.py
Real output:
TTFT (first visible token): 2.775s
Total time (last token): 7.584s
This is the entire lesson in two numbers. The total time barely changed (7.584s versus 7.428s before; the model is doing the same amount of work either way). But with streaming, the user sees the first word of the answer at 2.775 seconds instead of waiting the full 7.4 to 7.6 seconds for the whole thing. The server did not get any faster. Your client just stopped hiding the first 63 percent of the wait behind a blank screen.
Step 4: Prove Longer Context Costs More at the Front, Not the Back
Since TTFT is dominated by prefill, and prefill has to read every token in your prompt (system message, conversation history, and the new user message), a bigger prompt should mean a slower start, regardless of how short the question itself is. Test this directly with a short system prompt versus one padded with about 2,000 tokens of filler:
import time
import requests
OLLAMA_URL = "http://localhost:11434/api/chat"
MODEL = "qwen3.5:4b"
QUESTION = "In one sentence, what should I check first if a batch export job times out?"
def run(label, system_prompt):
payload = {
"model": MODEL,
"messages": [
{"role": "system", "content": system_prompt},
{"role": "user", "content": QUESTION},
],
"stream": False,
"think": False,
}
start = time.perf_counter()
data = requests.post(OLLAMA_URL, json=payload, timeout=120).json()
wall = time.perf_counter() - start
ns_to_s = 1_000_000_000
print(f"{label}: {data['prompt_eval_count']} tokens, "
f"prefill {data['prompt_eval_duration']/ns_to_s:.3f}s, "
f"wall {wall:.3f}s")
run("Short context", "You are a helpful assistant.")
run("Long context", "You are a helpful assistant. Prior history:\n" + "Ticket detail. " * 500)
Real output:
Short context: 40 tokens, prefill 0.538s, wall 4.911s
Long context: 2050 tokens, prefill 4.144s, wall 9.947s
The prompt grew 51.2 times larger (40 tokens to 2,050 tokens) and prefill duration grew 7.7 times (0.538s to 4.144s). It did not grow by the same 51x factor, because a GPU processes a big batch of prompt tokens more efficiently per token than a tiny one (this test’s prefill rate went from 74.4 tokens/sec on the short prompt to 494.7 tokens/sec on the long one), but the absolute time still went up substantially, and that time is paid before the model can say a single word. This is why agent frameworks that quietly accumulate conversation history, or that paste a large document into the system prompt on every turn, feel slower and slower to start responding as a session goes on, even though the question being asked each time is short.
Step 5: Measure the Cost of Sequential Tool Calls
TTFT and streaming fix how long the model takes to start talking. They do nothing for time your agent spends waiting on tools, like calling a weather API, a search engine, or a database, before it can even ask the model to respond. If an agent needs results from several independent tools and calls them one after another, it pays for the sum of every call’s latency.
Test this with a real, free, no-key weather API (Open-Meteo) called for four different cities:
import time
import requests
CITIES = {
"Berlin": (52.52, 13.41), "Tokyo": (35.68, 139.69),
"Sao Paulo": (-23.55, -46.63), "Nairobi": (-1.29, 36.82),
}
def fetch_weather(city):
lat, lon = CITIES[city]
url = (f"https://api.open-meteo.com/v1/forecast?latitude={lat}"
f"&longitude={lon}¤t=temperature_2m,wind_speed_10m&timezone=auto")
r = requests.get(url, timeout=30, headers={"User-Agent": "latency-demo"})
return r.json()["current"]
start = time.perf_counter()
for city in CITIES:
fetch_weather(city)
print(f"Sequential total: {time.perf_counter() - start:.3f}s")
Run it:
python 04_sequential.py
fetched Berlin in 0.604s
fetched Tokyo in 0.600s
fetched Sao Paulo in 0.600s
fetched Nairobi in 0.596s
Sequential total wall time: 2.400s
Each call takes roughly 0.6 seconds, and four calls in a row take roughly 2.4 seconds, the sum of all four. None of these four lookups depends on the result of another; there is no reason to wait for Berlin’s weather before asking for Tokyo’s.
Step 6: Parallelize Independent Tool Calls
Python’s concurrent.futures.ThreadPoolExecutor lets you issue multiple I/O-bound calls, like HTTP requests, at the same time and collect the results as they arrive. Since each of these calls spends almost all of its time waiting on the network rather than using the CPU, running them concurrently does not require anything fancier than threads:
import time
import concurrent.futures
import requests
def fetch_weather(city):
# same function as Step 5
...
start = time.perf_counter()
with concurrent.futures.ThreadPoolExecutor(max_workers=4) as pool:
futures = {pool.submit(fetch_weather, city): city for city in CITIES}
for fut in concurrent.futures.as_completed(futures):
fut.result()
print(f"Parallel total: {time.perf_counter() - start:.3f}s")
Berlin resolved at t=0.769s
Tokyo resolved at t=0.783s
Sao Paulo resolved at t=0.793s
Nairobi resolved at t=0.793s
Parallel total wall time: 0.794s
Four calls that took 2.400 seconds one after another took 0.794 seconds issued together, a 3.02x speedup. Notice the parallel run’s total (0.794s) is close to the slowest individual call, not the sum of all four; that is the general rule for parallelizing independent I/O-bound work. The more independent calls an agent needs to make, the bigger this gap gets. With four calls it is a 3x difference; with a dozen independent lookups, sequential execution would be more than 10x slower than parallel for no benefit at all.
Step 7: Put It Together: A Streaming Agent With Parallel Tool Calls
A real tool-using agent runs in three phases, and each phase needs a different one of the techniques above:
- Decision call (must be non-streaming). Ask the model, with tools available, what it wants to do. You cannot act on a partial tool call, so you need the complete
tool_callsobject before moving on. This call cannot usefully be streamed. - Tool execution (run independent calls in parallel). If the model asked for more than one tool call and they do not depend on each other, execute them concurrently, exactly as in Step 6.
- Final answer (stream it). Once tool results are back, ask the model to respond with the results in context, and stream that response so the user sees it start immediately.
Ollama’s tool format looks like this. You describe available tools as JSON schema:
TOOLS = [{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current temperature and wind speed for a city",
"parameters": {
"type": "object",
"properties": {"city": {"type": "string"}},
"required": ["city"],
},
},
}]
The model’s response includes a tool_calls array instead of a normal answer when it wants to use one:
"message": {
"role": "assistant", "content": "",
"tool_calls": [{"function": {"name": "get_weather", "arguments": {"city": "Berlin"}}}]
}
You send results back with "role": "tool":
messages.append({"role": "tool", "content": result_string, "tool_name": "get_weather"})
Putting the three phases together and asking “Compare the weather right now in Berlin and Tokyo. Which city is warmer?”, with two independent tool calls to resolve:
=== Step A: decision call (non-streaming) ===
Decision call took 6.535s
Model requested 2 tool call(s): Berlin, Tokyo
=== Step B: execute tool calls in parallel ===
Berlin -> 26.7 C, wind 14.4 km/h (t=0.746s)
Tokyo -> 27.1 C, wind 3.3 km/h (t=0.773s)
All tool calls resolved in 0.773s wall time
=== Step C: stream the final answer ===
Final-answer TTFT: 3.010s
Final-answer total: 10.801s
Answer: Based on the current weather data:
- Berlin: 26.7C (80F) with a wind speed of 14.4 km/h.
- Tokyo: 27.1C (81F) with a wind speed of 3.3 km/h.
Tokyo is currently warmer. While the difference in temperature between the
two cities is only minimal (0.4 degrees), Tokyo's current reading is
slightly higher than Berlin's.
The model correctly decided it needed both cities’ data and requested both tool calls in a single decision call, resolved them in parallel in 0.773 seconds instead of roughly 1.5 seconds sequentially, and streamed a correct final comparison starting 3.01 seconds after that call began. Total time to the last token was about 18 seconds (6.535 plus 0.773 plus 10.801), but the user saw the answer start at roughly 10.3 seconds in (6.535 plus 0.773 plus 3.010), not 18.
Common Mistakes and Gotchas
- Streaming does not make a slow model fast. It changes when content appears, not how long the underlying work takes. If your total generation time is the real problem, streaming will not fix it; shortening the requested output will.
- You cannot stream a tool-calling decision. The model needs to finish deciding which tools to call, and with what arguments, before you can act on that decision. Only the final answer, after tool results are available, is safe to stream in a simple client like the ones in this tutorial.
- Parallelizing only helps for truly independent calls. If one tool’s input depends on another tool’s output, they have to run in order no matter what. Check the dependency before reaching for a thread pool.
- Hidden reasoning tokens can dominate your timing without you noticing. As shown in Step 1, a reasoning model can generate over a thousand invisible tokens for a three-sentence visible answer. If your numbers look far slower than expected, compare
eval_countto the length of the visible response, and pass"think": falseif you do not need the reasoning shown. - Prefill cost is paid on every single call, not once. A bloated system prompt or an ever-growing conversation history slows down the start of every turn in a session, not just the first one, as demonstrated in Step 4.
How to Verify It All Works End to End
Re-run the Step 7 script and check that the reported numbers are internally consistent: the final-answer TTFT plus the two calls before it should be noticeably less than what fully sequential execution (decision call, then each tool call one after another, then the full unstreamed final answer) would have taken. In this tutorial’s run, streaming and parallelizing brought the user-perceived wait down to about 10.3 seconds against a true total of about 18 seconds, and against an even larger number if the two tool calls had been sequential instead of parallel. If your own numbers show the streamed TTFT roughly matching or exceeding the total time, something is wrong, most likely the final call was not actually sent with stream=True, or your loop is buffering all chunks before printing instead of displaying them as they arrive.
Next Steps
From here, try swapping the mock weather tool for a real tool your own agent needs, and confirm the same parallel-versus-sequential gap shows up. Watch what you accumulate into your system prompt and message history over a long-running session, using the technique from Step 4 to check how much prefill cost you are adding on every turn. If you want to see how caching interacts with prefill for repeated or extended conversations, this site’s prompt cache hit rate tutorial is a natural next read, and if you want to route different requests to different-sized models based on how complex they are, see the multi-tier LLM router tutorial. The same three habits practiced here, trimming what you send, overlapping independent waits, and surfacing output as soon as it exists rather than after everything finishes, keep paying off as an agent grows: they are usually cheaper to apply than swapping in a bigger model or renting more GPU capacity, and they are worth exhausting first.








No Comment! Be the first one.