How to Build a Local Tool-Calling AI Agent With LFM2.5-2.6B and Hugging Face Transformers
A hands-on walkthrough of running Liquid AI's edge-optimized LFM2.5-2.6B model locally with Hugging Face Transformers and building a real tool-calling agent that checks live weather and does math on...
Most local AI agent tutorials start with a multi-billion-parameter model and a beefy GPU. LFM2.5-2.6B, a small edge-optimized language model that Liquid AI released on August 4, 2026, takes the opposite approach: it is built to run agentic, tool-calling workloads on ordinary hardware, including a laptop CPU with no GPU at all. In this tutorial you will install it locally with Hugging Face Transformers, give it two real tools (a live weather lookup and a calculator), and build a small Python agent loop that lets the model decide when to call those tools, execute them, and use the results to answer questions it could not answer on its own.
Table Of Content
- What LFM2.5-2.6B Is, and Why It Is Different From a Typical Chat Model
- Prerequisites
- Step 1: Create an Isolated Python Environment
- Step 2: Install PyTorch, Transformers, and Accelerate
- Step 3: Load the Model and Run Your First Chat Completion
- What device_map=”auto” Actually Chose Here
- Step 4: Understand (and Correctly Strip) the Thinking Block
- Gotcha: There Is No Opening <think> Tag in the Generated Text
- Gotcha: Thinking Consumes Your Token Budget
- Step 5: Give the Model Real Tools
- Step 6: See What a Real Tool Call From the Model Looks Like
- Step 7: Safely Parse the Model’s Tool Call
- Step 8: Wrap It Into a Reusable Agent Loop
- Real Runs: Three Different Questions Through the Same Agent
- Query 1: A Question That Needs the Weather Tool
- Query 2: A Question That Needs the Calculator
- Query 3: A Question That Needs No Tool at All
- Step 9: Measure What You Are Actually Getting
- Common Mistakes and Gotchas
- How to Verify Everything Works End to End
- Next Steps
Everything in this post was run end to end on a real Windows sandbox with a 14-core CPU and no GPU acceleration used, so every command, timing number, and piece of output you will see below is a genuine capture, not a copy of Liquid AI’s own benchmark numbers.
What LFM2.5-2.6B Is, and Why It Is Different From a Typical Chat Model
A “language model” here just means software trained to predict the next chunk of text given everything that came before it. Most of the well-known ones, GPT-style or otherwise, are trained primarily to hold a conversation. LFM2.5-2.6B was trained differently: Liquid AI built it specifically to act as an agent, meaning a model that can decide, mid-conversation, that it needs to call an external function (a “tool”) to get information it does not already know, wait for the result, and then continue reasoning with that new information.
A few concrete facts about the model, all confirmed directly against Liquid AI’s own model card and blog post rather than taken on faith:
- Size: Liquid AI’s model card lists 2.69 billion parameters. This tutorial independently summed every loaded weight tensor rather than just quoting that figure, and got 2,697,198,592, which rounds to 2.70 billion, close enough to be the same model but a reminder that a vendor’s headline parameter count and a raw tensor sum do not always round to the identical second decimal.
- Architecture: 30 layers, a hybrid mix of 22 “double-gated short convolution” blocks and 8 grouped-query attention (GQA) blocks. This hybrid design is part of why the model can run fast on CPUs: convolution blocks are cheaper to evaluate than full attention.
- Context window: 131,072 tokens (128K), enough to hold long tool-result histories or long documents.
- Training: pretrained on roughly 34 trillion tokens, then post-trained in multiple stages (supervised fine-tuning weighted toward tool use, teacher-specialist distillation, and agentic reinforcement learning inside real agent harnesses) specifically to make it reliable at tool calling.
- License: the LFM Open License v1.0, from Liquid AI, Inc. It is free to use, including for commercial use, for any organization with under $10,000,000 in annual revenue, and unconditionally free for qualified non-profits. Organizations at or above that revenue threshold need a separate commercial license from Liquid AI. Read the full terms before using this model inside a large company.
Liquid AI explicitly recommends LFM2.5-2.6B for “agentic workloads, tool use, data extraction, RAG, and long-context workflows,” and explicitly does not recommend it for agentic coding or knowledge-heavy tasks, where a larger model will do better. Keep that scope in mind: this is a small, fast specialist, not a general-purpose replacement for a frontier model.
Prerequisites
You will need:
- A computer running Windows, macOS, or Linux with Python 3.10 or newer installed. This tutorial was run on Python 3.13.
- About 6 GB of free disk space: roughly 5.1 GB for the model weights themselves, plus about 900 MB for PyTorch and its dependencies (PyTorch’s CPU build alone accounts for most of that, at around 530 MB installed).
- An internet connection to download the model the first time. After that, everything in this tutorial runs fully offline except the weather tool, which calls a free public API.
- No GPU is required. Every test in this tutorial ran on CPU only. If you do have an NVIDIA GPU, see the note near the end about why a small (4 GB class) card will not necessarily help here.
- Basic familiarity with Python: functions, dictionaries, and running scripts from a terminal. You do not need any prior experience with Hugging Face Transformers or with AI agents specifically; both are explained as we go.
- Patience. CPU inference on a 2.6 billion parameter model is not instant. Expect single-digit tokens per second unless you have a very fast chip; this tutorial measures the real number on its own test machine so you know what to expect on yours.
Step 1: Create an Isolated Python Environment
Before installing anything, create a dedicated virtual environment so this project’s dependencies (a specific PyTorch and Transformers version) cannot collide with anything else on your system. Open a terminal, create a project folder, and run:
python -m venv venv
On Windows, activate it with:
venv\Scripts\activate
On macOS or Linux, activate it with:
source venv/bin/activate
You will know it worked because your terminal prompt now shows (venv) at the start of the line. Everything you install from this point forward will be isolated to this folder, and you can delete the whole thing later with no side effects on the rest of your system.
Step 2: Install PyTorch, Transformers, and Accelerate
With the virtual environment active, install the three packages this tutorial needs:
pip install torch transformers accelerate requests
Here is what each one does, since it matters for understanding errors later:
- torch (PyTorch) is the numerical library that actually runs the model’s matrix multiplications. If you do not have an NVIDIA GPU, or you have one with limited memory like the 4 GB card in this tutorial’s test machine, plain
pip install torchon a machine without visible CUDA hardware installs a CPU-only build automatically. This tutorial ended up ontorch 2.13.0+cpu, confirmed by printingtorch.__version__. - transformers is Hugging Face’s library for loading and running pretrained models like LFM2.5-2.6B from a simple model name, instead of hand-writing the architecture yourself. This tutorial used
transformers 5.14.1. - accelerate lets Transformers automatically decide which device (CPU, GPU, or a mix) to place each part of the model on via
device_map="auto". - requests is only used later, for the live weather tool, to make an HTTP call to a public weather API.
Note the exact versions above. Hugging Face Transformers moves fast, and as you will see in Step 4, at least one behavior in this tutorial changed between major Transformers versions.
Step 3: Load the Model and Run Your First Chat Completion
Create a file named chat_test.py with the following:
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "LiquidAI/LFM2.5-2.6B"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(
model_id,
device_map="auto",
dtype="bfloat16",
)
messages = [{"role": "user", "content": "What is C. elegans?"}]
inputs = tokenizer.apply_chat_template(
messages,
add_generation_prompt=True,
return_tensors="pt",
tokenize=True,
return_dict=True,
).to(model.device)
output = model.generate(
**inputs,
do_sample=True,
temperature=0.1,
top_k=50,
repetition_penalty=1.1,
max_new_tokens=900,
)
new_tokens = output[0][inputs["input_ids"].shape[-1]:]
print(tokenizer.decode(new_tokens, skip_special_tokens=False))
The temperature=0.1, top_k=50, and repetition_penalty=1.1 values are not arbitrary; they are the exact generation parameters Liquid AI publishes on the model’s own card as the recommended settings, so this tutorial uses those rather than inventing its own.
Run it:
python chat_test.py
The first run downloads the model from Hugging Face, two safetensors shard files totaling about 5.1 GB on disk (confirmed by checking the cache folder size directly). On this tutorial’s test machine, with a normal broadband connection, the download and first load together took about one minute (3.7 seconds to load the tokenizer, 57 seconds to fetch and load the weights); every run after that reuses Hugging Face’s local cache (~/.cache/huggingface) and loads in under a second, because the on-disk weights are memory-mapped straight into the process rather than re-read and re-parsed.
What device_map=”auto” Actually Chose Here
This tutorial’s test machine does have a discrete NVIDIA GPU, but only 4 GB of video memory, of which about 3.6 GB was free. LFM2.5-2.6B’s weights alone need roughly 5.4 GB in bfloat16 (2.69 billion parameters times 2 bytes each), which does not fit in 3.6 GB free VRAM. Rather than fight that constraint, this tutorial pinned everything to CPU explicitly with device_map="cpu" for every timed test below, which is also the realistic path for most readers who do not have a GPU at all. If you do have a GPU with at least 8 to 10 GB free, device_map="auto" will place the model there instead and generation will be dramatically faster; this tutorial did not have that hardware available to verify GPU numbers firsthand, so it only reports what it actually measured, which is the CPU path.
Step 4: Understand (and Correctly Strip) the Thinking Block
LFM2.5-2.6B is what Liquid AI calls a “pure reasoning model”: before it writes a final answer, it always writes out a scratchpad of its reasoning first, wrapped in a <think> block. Run the script above and you will see output that looks like this (real, unedited capture from this tutorial’s test machine, for the prompt “What is C. elegans? Answer in 2 sentences.”):
The user wants to know what *C. elegans* is, and the constraint is that
it must be answered in exactly two sentences.
1. **Identify the subject:** *C. elegans* (Caenorhabditis elegans).
2. **Determine key characteristics:**
* It's a nematode worm.
* It has a fixed number of cells (programmed cell death).
* It's a model organism for biology/genetics.
...
Check sentence count:
1. *C. elegans* is a microscopic roundworm species that serves as a
fundamental model organism in biology due to its simple anatomy and
short life cycle.
2. It is renowned for having a fixed number of cells throughout its
life, which makes it ideal for studying developmental biology and
genetic regulation.
Total: 2 sentences. Perfect.</think>*C. elegans* is a microscopic
roundworm species that serves as a fundamental model organism in
biology due to its simple anatomy and short life cycle. It is renowned
for having a fixed number of cells throughout its life, which makes it
ideal for studying developmental biology and genetic regulation.
Two things about that output are easy to get wrong, and this tutorial got the first one wrong on the first attempt before catching it.
Gotcha: There Is No Opening <think> Tag in the Generated Text
Look closely at the raw output above: it ends with ...Total: 2 sentences. Perfect.</think>, a closing tag, but there is no <think> opening tag anywhere in it. That is not a bug in the model. Liquid AI’s chat template injects the literal text <think> at the end of the prompt, right after the model’s turn begins, specifically so the model always starts reasoning immediately. Since that opening tag lives in the prompt, decoding only the newly generated tokens (as this tutorial’s code does with output[0][inputs["input_ids"].shape[-1]:]) will never show it; only the model-generated closing </think> appears.
This tutorial’s first attempt at a cleanup function used a regular expression that looked for a matching <think>...</think> pair inside the generated text, and it silently matched nothing, because the opening half of that pair was never there to match. The fix is to stop looking for a pair and instead split on the closing tag alone:
def strip_think(raw_text: str) -> str:
# The opening tag lives in the PROMPT, not the generated
# text, so only ever shows up in what you decode here.
cleaned = raw_text.split("", 1)[-1]
for special in ("<|im_end|>", "<|im_start|>"):
cleaned = cleaned.replace(special, "")
return cleaned.strip()
Verify it worked by printing both the raw and cleaned versions side by side; the cleaned version should contain only the two-sentence answer, with no reasoning trace and no leftover special tokens.
Gotcha: Thinking Consumes Your Token Budget
The second issue is more consequential. This tutorial’s very first attempt used max_new_tokens=400, a number that sounds generous for a two-sentence answer. It was not. The model spent the entire 400-token budget drafting and second-guessing its reasoning and never emitted a final answer at all, the generation simply stopped mid-thought when it hit the token limit. Raising the limit to max_new_tokens=900 was enough for the example above, which used 486 tokens total, but a more complex question, or one that also needs a tool call and a follow-up response, can need more. Budget generously, and if you need a hard ceiling on response time, measure it empirically for your own prompts rather than guessing.
Step 5: Give the Model Real Tools
“Tool calling” (also called “function calling”) means giving a language model a menu of Python functions it is allowed to request, described in a structured format it can read, without giving it the ability to run arbitrary code. The model itself never executes anything; it only ever outputs text saying “call this function with these arguments,” and your own code decides whether, and how, to actually run it.
Create a file named tools.py with two real, working tools: a live weather lookup using Open-Meteo’s free API (no API key required) and a calculator that evaluates arithmetic without using Python’s dangerous eval():
import ast
import operator
import requests
GEOCODE_URL = "https://geocoding-api.open-meteo.com/v1/search"
FORECAST_URL = "https://api.open-meteo.com/v1/forecast"
WEATHER_CODES = {
0: "clear sky", 1: "mainly clear", 2: "partly cloudy", 3: "overcast",
45: "fog", 61: "slight rain", 63: "moderate rain", 65: "heavy rain",
71: "slight snow", 80: "slight rain showers", 95: "thunderstorm",
}
def get_weather(city: str) -> dict:
"""Look up current weather for a city using Open-Meteo's free API."""
geo = requests.get(GEOCODE_URL, params={"name": city, "count": 1}, timeout=10).json()
results = geo.get("results") or []
if not results:
return {"error": f"could not find a location named {city!r}"}
place = results[0]
fc = requests.get(
FORECAST_URL,
params={
"latitude": place["latitude"],
"longitude": place["longitude"],
"current": "temperature_2m,weather_code",
},
timeout=10,
).json()
current = fc.get("current", {})
code = current.get("weather_code")
return {
"city": place.get("name"),
"country": place.get("country"),
"temperature_c": current.get("temperature_2m"),
"condition": WEATHER_CODES.get(code, f"code {code}"),
}
_ALLOWED_BINOPS = {
ast.Add: operator.add, ast.Sub: operator.sub, ast.Mult: operator.mul,
ast.Div: operator.truediv, ast.Pow: operator.pow, ast.Mod: operator.mod,
}
def _safe_eval(node):
if isinstance(node, ast.Constant) and isinstance(node.value, (int, float)):
return node.value
if isinstance(node, ast.BinOp) and type(node.op) in _ALLOWED_BINOPS:
return _ALLOWED_BINOPS[type(node.op)](_safe_eval(node.left), _safe_eval(node.right))
if isinstance(node, ast.UnaryOp) and isinstance(node.op, ast.USub):
return -_safe_eval(node.operand)
raise ValueError(f"disallowed expression: {ast.dump(node)}")
def calculate(expression: str) -> dict:
"""Evaluate a numeric arithmetic expression without calling eval()."""
try:
tree = ast.parse(expression, mode="eval")
return {"expression": expression, "result": _safe_eval(tree.body)}
except Exception as exc:
return {"expression": expression, "error": str(exc)}
The calculator deliberately avoids eval(). Python’s ast module lets you parse an expression into a tree and walk it yourself, checking every node against an explicit allowlist of safe operations, so an input like "10 / 0" raises a normal, catchable ZeroDivisionError that this tutorial’s code turns into a clean error dictionary, while something like "__import__('os').system('...')" is rejected outright, since ast.parse in mode="eval" only accepts a single expression and the walker only recognizes numeric literals and arithmetic operators, nothing else.
Test both functions on their own before wiring up the model, so you know any bugs you hit later are in the agent logic, not the tools themselves. Real output from this tutorial’s test machine:
>>> from tools import get_weather, calculate
>>> get_weather("Reykjavik")
{'city': 'Reykjavik', 'country': 'Iceland', 'temperature_c': 14.1, 'condition': 'clear sky'}
>>> calculate("(19 + 5) * 3 - 7")
{'expression': '(19 + 5) * 3 - 7', 'result': 65}
>>> calculate("10 / 0")
{'expression': '10 / 0', 'error': 'division by zero'}
Now add the JSON schema descriptions that tell the model what each tool does and what arguments it takes. This is the same style OpenAI, Anthropic, and most other tool-calling APIs use, so it will look familiar if you have used any of them:
TOOL_SCHEMAS = [
{
"type": "function",
"function": {
"name": "get_weather",
"description": "Get the current weather for a named city.",
"parameters": {
"type": "object",
"properties": {
"city": {"type": "string", "description": "City name, e.g. 'Tokyo'"},
},
"required": ["city"],
},
},
},
{
"type": "function",
"function": {
"name": "calculate",
"description": "Evaluate a numeric arithmetic expression, e.g. '(19 + 5) * 3'.",
"parameters": {
"type": "object",
"properties": {
"expression": {"type": "string", "description": "A Python-style arithmetic expression"},
},
"required": ["expression"],
},
},
},
]
TOOL_FUNCTIONS = {"get_weather": get_weather, "calculate": calculate}
Step 6: See What a Real Tool Call From the Model Looks Like
Pass tools=TOOL_SCHEMAS into apply_chat_template(), and ask a question the model cannot answer from its own training data, like today’s weather somewhere:
from tools import TOOL_SCHEMAS
messages = [{"role": "user", "content": "What's the weather like in Reykjavik right now?"}]
inputs = tokenizer.apply_chat_template(
messages, tools=TOOL_SCHEMAS, add_generation_prompt=True,
return_tensors="pt", tokenize=True, return_dict=True,
).to(model.device)
output = model.generate(**inputs, do_sample=True, temperature=0.1, top_k=50,
repetition_penalty=1.1, max_new_tokens=900)
print(tokenizer.decode(output[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=False))
Real, unedited output from this tutorial’s test machine:
The user is asking about the current weather in Reykjavik. I have
access to a `get_weather` function that can provide this information.
The function requires a city parameter, which should be "Reykjavik"
based on the user's question.
Let me call this function to get the current weather for
Reykjavik.</think><|tool_call_start|>[get_weather(city='Reykjavik')]<|tool_call_end|><|im_end|>
This is worth reading carefully, because it is different from how most tool-calling models behave. Instead of emitting a JSON object like {"name": "get_weather", "arguments": {"city": "Reykjavik"}}, LFM2.5-2.6B writes an actual Python-looking function call, wrapped between two special tokens: <|tool_call_start|> and <|tool_call_end|>. Liquid AI calls this its “Pythonic” tool call format. It matters practically: if you try to hand that string straight to json.loads(), expecting the JSON-style output other model families use, it will fail immediately, because get_weather(city='Reykjavik') is not valid JSON.
Step 7: Safely Parse the Model’s Tool Call
Since the tool call is a real Python call expression, the safe way to parse it is the same technique used for the calculator: Python’s ast module, never eval(). Add this to a new file, agent.py:
import ast
import re
TOOL_CALL_RE = re.compile(r"<\|tool_call_start\|>(.*?)<\|tool_call_end\|>", re.DOTALL)
def parse_tool_calls(raw_text: str):
"""Parse LFM2.5's Pythonic tool-call syntax into a list of
{"name": ..., "arguments": {...}} dicts. Uses ast, never eval()."""
match = TOOL_CALL_RE.search(raw_text)
if not match:
return []
tree = ast.parse(match.group(1).strip(), mode="eval")
if not isinstance(tree.body, ast.List):
raise ValueError(f"expected a list of calls, got: {ast.dump(tree.body)}")
calls = []
for element in tree.body.elts:
if not isinstance(element, ast.Call):
raise ValueError(f"expected a function call, got: {ast.dump(element)}")
calls.append({
"name": element.func.id,
"arguments": {kw.arg: ast.literal_eval(kw.value) for kw in element.keywords},
})
return calls
Three things this parser deliberately does, each for a reason:
- It uses
ast.parse(..., mode="eval"), which only accepts a single Python expression, never a full statement, assignment, or import. That alone rules out most injection attempts before any further checks run. - It walks the parsed tree and only accepts a top-level list of
ast.Callnodes, matching the exact shape the model actually produces ([func(arg=val)]). Anything else raises immediately instead of guessing. - It reads each argument value with
ast.literal_eval(), which only ever produces plain Python literals (strings, numbers, lists, dicts, booleans,None). It cannot call a function, access an attribute, or execute anything, unlikeeval().
Now execute the parsed call for real and feed the result back to the model. Liquid AI’s own documentation shows tool results going back in as a message with role="tool", whose content is the JSON-encoded result. Continue the same conversation from Step 6:
from tools import TOOL_FUNCTIONS
import json
raw = tokenizer.decode(output[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=False)
calls = parse_tool_calls(raw)
print(calls)
# -> [{'name': 'get_weather', 'arguments': {'city': 'Reykjavik'}}]
call = calls[0]
result = TOOL_FUNCTIONS[call["name"]](**call["arguments"])
print(result)
# -> {'city': 'Reykjavik', 'country': 'Iceland', 'temperature_c': 14.1, 'condition': 'clear sky'}
messages.append({
"role": "assistant",
"tool_calls": [{"function": {"name": call["name"], "arguments": call["arguments"]}}],
})
messages.append({"role": "tool", "content": json.dumps([result])})
inputs2 = tokenizer.apply_chat_template(
messages, tools=TOOL_SCHEMAS, add_generation_prompt=True,
return_tensors="pt", tokenize=True, return_dict=True,
).to(model.device)
output2 = model.generate(**inputs2, do_sample=True, temperature=0.1, top_k=50,
repetition_penalty=1.1, max_new_tokens=900)
final = tokenizer.decode(output2[0][inputs2["input_ids"].shape[-1]:], skip_special_tokens=False)
print(final)
Real output from the second call, after the tool result was fed back in:
The user is asking about the current weather in Reykjavik. I have
already called the get_weather function and received the response.
The data shows:
- City: Reykjavik
- Country: Iceland
- Temperature: 14.1°C
- Condition: clear sky
I should provide this information to the user in a friendly and
helpful way.</think>The current weather in Reykjavik, Iceland is
**clear sky** with a temperature of **14.1°C**. It's a nice, mild
day there!
The final sentence is the model’s own words, generated from the real tool result, not a template this tutorial wrote. The temperature and condition in that sentence match exactly what get_weather() actually returned two steps earlier.
Passing the assistant’s request as a structured tool_calls field, rather than replaying its raw <|tool_call_start|>...<|tool_call_end|> text verbatim, matters: the chat template itself knows how to re-render a structured tool call back into the model’s expected Pythonic format on the next turn, so you never need to hand-construct that special-token syntax yourself.
Step 8: Wrap It Into a Reusable Agent Loop
Put the pieces from Steps 6 and 7 together into a function that keeps calling tools, feeding back results, and re-prompting the model until it gives a final plain-text answer or hits a safety limit on the number of rounds. This is the core structure behind most production tool-using agents, regardless of which model or framework sits underneath:
import json
import time
from tools import TOOL_FUNCTIONS, TOOL_SCHEMAS
def strip_think(raw_text: str) -> str:
cleaned = raw_text.split("", 1)[-1]
for special in ("<|im_end|>", "<|im_start|>"):
cleaned = cleaned.replace(special, "")
return cleaned.strip()
def run_agent(tokenizer, model, user_prompt, max_new_tokens=900, max_tool_rounds=3):
messages = [{"role": "user", "content": user_prompt}]
for round_num in range(max_tool_rounds + 1):
inputs = tokenizer.apply_chat_template(
messages, tools=TOOL_SCHEMAS, add_generation_prompt=True,
return_tensors="pt", tokenize=True, return_dict=True,
).to(model.device)
output = model.generate(
**inputs, do_sample=True, temperature=0.1, top_k=50,
repetition_penalty=1.1, max_new_tokens=max_new_tokens,
)
raw = tokenizer.decode(output[0][inputs["input_ids"].shape[-1]:], skip_special_tokens=False)
calls = parse_tool_calls(raw)
if not calls:
return strip_think(raw)
messages.append({
"role": "assistant",
"tool_calls": [{"function": {"name": c["name"], "arguments": c["arguments"]}} for c in calls],
})
for call in calls:
fn = TOOL_FUNCTIONS.get(call["name"])
result = fn(**call["arguments"]) if fn else {"error": f"unknown tool {call['name']}"}
messages.append({"role": "tool", "content": json.dumps([result])})
return "[gave up after max tool rounds]"
Notice that the loop calls parse_tool_calls() on every round and only stops once it comes back empty, meaning the model chose to answer in plain text instead of requesting another tool. This lets the same function handle a question that needs zero tools, one tool, or a chain of several, without any special-casing.
Real Runs: Three Different Questions Through the Same Agent
This tutorial ran the exact run_agent() function above against three different prompts on the same loaded model, to confirm it correctly decides when to call a tool and when not to.
Query 1: A Question That Needs the Weather Tool
Prompt: “What’s the weather like in Reykjavik right now?” This is the same exchange walked through step by step above. The agent took 2 rounds: round 0 produced the get_weather call (17.4 seconds), the tool ran and returned a real result, then round 1 produced the final answer (21.6 seconds). Total real-world time for this one question: about 39 seconds on CPU, most of it spent thinking and generating, not on the near-instant weather API call itself.
Query 2: A Question That Needs the Calculator
Prompt: “What is (19 + 5) * 3, minus 7?” Real captured round 0 output:
The user is asking me to calculate (19 + 5) * 3 - 7. This is a simple
arithmetic expression that I can solve using the calculate tool.
Let me break it down:
1. First, calculate (19 + 5) = 24
2. Then multiply by 3: 24 * 3 = 72
3. Finally subtract 7: 72 - 7 = 65
I should use the calculate function to get the precise result.</think><|tool_call_start|>[calculate(expression='(19 + 5) * 3 - 7')]<|tool_call_end|><|im_end|>
Note that the model worked out the answer (65) inside its own reasoning before ever calling the tool, then called calculate() anyway to get a precise, verifiable result rather than trusting its own mental arithmetic. The tool executed and returned {'expression': '(19 + 5) * 3 - 7', 'result': 65}, matching its own draft calculation, and round 1 produced the final answer: The result of (19 + 5) * 3 - 7 is 65. Two rounds, 22.3 and 20.1 seconds.
Query 3: A Question That Needs No Tool at All
Prompt: “In one sentence, what programming language is Django written in?” This tests the opposite behavior: does the agent correctly recognize when neither registered tool applies, rather than forcing a tool call it does not need? Real captured output, in a single round (29.8 seconds, no tool round-trip needed):
The user is asking about what programming language Django is written
in. This is a factual question that I should answer directly. I know
from my training data that Django is written in Python. However, let
me verify this using the available tools if possible.
Looking at the available tools:
1. `get_weather` - for weather information
2. `calculate` - for arithmetic expressions
Neither of these tools can help answer a factual question about
programming languages. The question is straightforward and doesn't
require any calculation or weather information. I should provide the
correct answer based on my knowledge.
Django is indeed written in Python. It's an open-source web framework
created by Kenneth Reitz and first released in 2005. So I can answer
this directly without needing to call any tools.</think>Django is
written in Python.
parse_tool_calls() correctly returned an empty list here, since no <|tool_call_start|> token appears anywhere in that output, so run_agent() returned the cleaned answer on round 0 with no tool round-trip at all.
This transcript is also a useful, honest warning about trusting a model’s hidden reasoning: the <think> block above misattributes Django’s creation to Kenneth Reitz, a real developer, but the one actually known for the unrelated requests and pipenv Python libraries. Django was actually created by Adrian Holovaty and Simon Willison at the Lawrence Journal-World newspaper, according to the Django project’s own FAQ. The release year, 2005, happens to be correct. The model’s own final, user-facing answer wisely stayed narrow, just “Django is written in Python.”, and never repeated the incorrect name. That is a real, unprompted example of why you should treat a model’s reasoning trace as scratch work, not as a verified fact source, even when the final answer it produces turns out fine.
Step 9: Measure What You Are Actually Getting
Benchmark numbers from a vendor’s own blog post are measured on their hardware, often with an optimized runtime. To know what your own setup will actually feel like, measure it yourself. This tutorial ran a short, fixed prompt through the loaded CPU model and timed only the generate() call:
import time
messages = [{"role": "user", "content": "Count from 1 to 5."}]
inputs = tokenizer.apply_chat_template(
messages, add_generation_prompt=True, return_tensors="pt",
tokenize=True, return_dict=True,
).to(model.device)
t0 = time.time()
output = model.generate(**inputs, do_sample=True, temperature=0.1, top_k=50,
repetition_penalty=1.1, max_new_tokens=300)
elapsed = time.time() - t0
new_tokens = output.shape[-1] - inputs["input_ids"].shape[-1]
print(f"{new_tokens} tokens in {elapsed:.2f}s -> {new_tokens/elapsed:.2f} tok/s")
Real output on this tutorial’s 14-core Intel Core Ultra 5 235, CPU only, plain PyTorch bfloat16, no quantization: 67 tokens in 10.65 seconds, 6.29 tokens per second. A separate, longer run earlier in this tutorial (486 tokens for the “2 sentences” answer in Step 4) measured 6.44 tokens per second, so this machine sustains a consistent 6.3 to 6.4 tokens per second across different prompts.
For comparison, Liquid AI’s own blog post reports 220 tokens per second on an Apple M5 Max and 113 tokens per second on an AMD Ryzen AI Max+ 395, both described as CPU inference, and both dramatically faster than what this tutorial measured with plain PyTorch on a 14-core Intel Core Ultra 5 235. The likely explanation, based on what Liquid AI documents elsewhere on the model card, is that those figures use a quantized, CPU-optimized runtime such as the GGUF build through llama.cpp, not the full-precision bfloat16 weights loaded directly through plain Transformers, which is what this tutorial deliberately used throughout to keep the code as close to Liquid AI’s own documented “Get Started” example as possible. If raw CPU speed matters for your use case, the GGUF version linked from the model’s Hugging Face page is the more appropriate starting point than the path in this tutorial.
Memory told a similar story. Liquid AI’s blog states the model runs in “under 2.5 GB of memory.” This tutorial measured actual process memory with psutil before and after generation and saw resident memory jump from under 0.4 GB (weights are memory-mapped from disk and mostly not yet touched right after loading) to about 5.3 GB once generation actually ran every layer, consistent with 2.69 billion parameters at 2 bytes each in bfloat16, plus overhead. The gap again points to a quantized runtime behind Liquid AI’s own number rather than a discrepancy in this tutorial’s measurement; if you need the model to fit in a smaller memory budget, look at the GGUF or ONNX builds rather than the raw safetensors checkpoint this tutorial uses.
Common Mistakes and Gotchas
Beyond the two thinking-block issues already covered in Step 4, watch for these:
apply_chat_template’s return type depends on your Transformers version. On transformers 5.14.1, calling tokenizer.apply_chat_template(..., return_tensors="pt", tokenize=True) without an explicit return_dict argument returned a BatchEncoding object, a dictionary-like container with input_ids and attention_mask keys, not a bare tensor. Code written against older examples that do input_ids = tokenizer.apply_chat_template(...) and then immediately call input_ids.shape will crash with a confusing AttributeError, because a BatchEncoding has no .shape. This tutorial hit that exact error on its first run, in a script that printed input_ids.shape[-1] right after calling apply_chat_template(). The real traceback, trimmed to the relevant frames:
KeyError: 'shape'
During handling of the above exception, another exception occurred:
Traceback (most recent call last):
File "...tokenization_utils_base.py", line 291, in __getattr__
raise AttributeError
AttributeError
The confusing part is that top-level AttributeError with no message. It happens because BatchEncoding.__getattr__ first tries to look up "shape" as a dictionary key (self.data["shape"]), fails with an internal KeyError, and re-raises that as a bare AttributeError instead. If you ever see this exact pattern, an AttributeError with no message that traces back through tokenization_utils_base.py, suspect a BatchEncoding where you expected a tensor before looking anywhere else.
The fix used throughout this tutorial is to pass return_dict=True explicitly and unpack the result into model.generate() with **inputs, which works whether or not your installed version defaults to a dict.
Forgetting the tools= argument. If you call apply_chat_template() without tools=TOOL_SCHEMAS, the model has no way to know a get_weather function exists, and it will simply answer from its own (incorrect, potentially outdated) beliefs instead of calling anything. It will not warn you; it will just guess.
The Pythonic tool call format is not JSON. If you have used OpenAI’s or Anthropic’s tool-calling APIs before, your instinct will be to reach for json.loads() on the model’s tool call output. That will raise a JSONDecodeError here, because LFM2.5-2.6B emits real Python call syntax like get_weather(city='Reykjavik'), not a JSON object.
How to Verify Everything Works End to End
Before relying on this setup for anything real, confirm each of the following, in order:
First, confirm the base model answers a plain question with no tools involved, and that the cleaned (post strip_think) output contains no leftover <think>, </think>, or <|im_end|> fragments. Second, confirm parse_tool_calls() correctly returns an empty list for a question that needs no tool, so your agent does not loop or hallucinate a tool call where none belongs; Query 3 above is exactly this check. Third, confirm a genuine tool call round-trips correctly: the model requests a tool, your code executes the real function (not a stub), the result is fed back with role="tool", and the model’s final answer actually reflects the tool’s real output rather than a guess, for example a temperature value in the final answer that matches what get_weather() actually returned. Fourth, deliberately test a failure path, such as asking about a nonexistent city or dividing by zero, and confirm your tool functions return a clean error dictionary instead of raising an uncaught exception that would crash the whole agent loop.
Next Steps
From here, a few directions are worth exploring. Try the LFM2.5-2.6B-GGUF build through llama.cpp if CPU speed matters more than staying inside the Transformers/PyTorch ecosystem; based on Liquid AI’s own published numbers, it should be substantially faster than the plain bfloat16 path this tutorial used. If you have an Apple Silicon Mac, the LFM2.5-2.6B-MLX build is built specifically for that hardware. For production serving rather than a single local script, Liquid AI lists day-one support for vLLM and SGLang, both of which handle batching many simultaneous requests far better than the single-request model.generate() calls used throughout this tutorial. And if two tools was not enough, the agent loop built in Step 8 does not care how many tools you register; add more entries to TOOL_SCHEMAS and TOOL_FUNCTIONS and the same loop will handle them.








No Comment! Be the first one.