TRENDING
Rows of identical brass-colored apartment mailboxes with small locks and name labels along an orange corridor wall
October 9, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
Street-level upward view of the Monetary Authority of Singapore building and neighbouring office towers under a pale sky
October 9, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
Cast-iron late Qing dynasty coin minting press with a large flywheel, displayed in a museum case
October 9, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google
Rows of closed oak library card catalog drawers, each with a brass pull and a blank label holder
October 9, 2026
How to Encrypt PII in Python and Keep It Searchable With Blind Indexes
Close-up of a vintage Western Electric manual telephone switchboard with orange lamps, red patch cords plugged into jacks, a rotary dial and a black handset
October 9, 2026
Microsoft’s Agent Lightning v1.0 Turns Agent Training Into a Sample-Accounting Problem
09 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Two orange safety relief valves on grey pressure vessels in an industrial plant
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
Yellow diamond-shaped merging traffic warning sign showing a side road joining a main road
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A lugworm lying on wet sand and mud at low tide
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 232 Posts
News 234 Posts
Learning Hub 204 Posts
Home/Articles/Microsoft’s Agent Lightning v1.0 Turns Agent Training Into a Sample-Accounting Problem
Articles

Microsoft’s Agent Lightning v1.0 Turns Agent Training Into a Sample-Accounting Problem

Microsoft’s rebuilt Agent Lightning trains agents through an unmodified harness and a recording proxy, which moves the hard part of agent reinforcement learning from rewriting the agent to counting...

October 7, 2026 16 Min Read
19

On October 7, Microsoft Research published a post on Agent Lightning v1.0, an open-source framework (MIT license, about 18,600 GitHub stars when I checked) that trains an AI agent with reinforcement learning while the agent keeps running inside the harness it ships with. A harness is the program that supplies an agent’s prompts, tools and control loop; Claude Code, Codex and OpenHands are examples. The post’s summary box makes four claims: a training paradigm in which the deployed harness takes part directly in training, a control plane in “roughly 3,500 lines of code,” agents that run as standard Kubernetes jobs, and a coding-agent recipe that raised Qwen3.5-9B from 41.8% to 56.4% on SWE-bench Verified, “using only about 6,000 training samples.”

Table Of Content

  • How the proxy changes what a trainer sees
  • The four counting problems the report names
  • Retokenization decides how many samples one attempt becomes
  • Advantage and loss: the fragmentation tax
  • What the 3,500 lines count
  • What the evidence supports
  • Six thousand samples, and the compute behind them
  • Reading 41.8% to 56.4%
  • The gateway is also the trust boundary
  • Authentication is opt-in
  • The agent’s shell inherits the key
  • Some guardrails live inside the agent’s reach
  • What I would do before pointing this at real workloads
  • The bottom line
  • Reproduce the two checks
  • Count the lines
  • Download and extract
  • The script
  • Test the chat template and the token drift

I read the post, the technical report (arXiv 2608.17528) and the v1.0.2 source, and ran two checks of my own. What ties them together is that the proxy that makes a harness easy to connect also hides everything the harness does between model calls, so the trainer must rebuild training samples from a stream of requests and responses it did not orchestrate. The report says existing frameworks “generally leave these issues underspecified.” That puts the counting rules at the center: how many samples one attempt becomes, how its reward is shared among them, what the headline numbers rest on, and who can reach the gateway that records everything.

How the proxy changes what a trainer sees

In reinforcement learning (RL) for agents, a model attempts a task, a checker scores the attempt, and the model is nudged toward behavior that scored well. One attempt is a rollout and its score is the reward; for a coding agent the reward is usually whether hidden tests pass after the agent edits a repository. The report says earlier RL frameworks, including verl, AReaL and slime, “generally require users to implement the agent loop directly inside the training framework,” which means rebuilding each harness inside the trainer. The blog post adds that “the rebuilt agent may no longer behave in the same way as the deployed agent.”

Agent Lightning’s answer, first described in an August 2025 paper and rebuilt for v1.0, is a recording proxy between the harness and the model. The harness keeps calling an OpenAI-compatible endpoint; the endpoint now belongs to Agent Lightning, which forwards each call to the model being trained and stores the prompt token IDs, response token IDs and log-probabilities. In the blog’s words, “simply point the endpoint that previously called the model API at Agent Lightning, and the training framework can observe and record its model calls.” The report calls the arrangement harnessed agentic RL and says verl Uni-Agent, AReaL 2.0, slime v0.3.0 and Polar have since adopted the same proxy approach, so most of what follows is not specific to this codebase. For harnesses as a category, see this site’s piece on DeepSeek’s harness release.

The repository’s README describes three components. The trainer runs verl and vLLM, builds training samples and updates the policy. The API gateway proxies model requests and captures training data. The rollout controller runs agents locally or as Kubernetes Jobs. The report’s Table 1 lists the gateway’s endpoints, among them routes to create and patch rollouts, append and read events, register and delete model endpoints, and forward a chat-completion call. The trainer sees only what the gateway recorded, which is the source of the first problem below.

The four counting problems the report names

The report lists retokenization and sample merging, advantage calculation, loss normalization and training-backend scheduling. The first three can be checked with a short script or with arithmetic, so I did.

Retokenization decides how many samples one attempt becomes

RL training needs the exact token IDs the model sampled, but a harness stores its conversation as text and sends the whole text back on the next call. If re-tokenizing that text does not reproduce the earlier tokens, the trainer cannot treat two consecutive calls as one continuous sequence. The report finds three causes: chat templates that render a full history differently from the sum of its parts (Qwen’s template can drop an earlier think marker, it says), decoding sampled tokens to text and tokenizing them again, and tool-call handlers that repair or reserialize a response before returning it.

Frameworks handle the mismatch in different ways. AReaL and verl Uni-Agent keep a buffer in the proxy and swap the buffered token IDs into the next prompt; the report objects that this “becomes off-policy stitching when it changes the prompt actually consumed during rollout.” Agent Lightning merges two calls only when the exact token prefix matches. Its trainer docs say: “If exact token-prefix continuity is broken, the aggregator starts a new training row instead of merging incompatible calls.”

To see the template effect, I rendered a five-call toy rollout through the Qwen3.5-9B tokenizer and chat template (transformers 5.19.0, Python 3.13.14). It is shaped like the shipped SWE-smith example: a system prompt, a task, then alternating assistant replies and observations that the harness sends back as user messages. For each call I asked whether the previous prompt plus the response the model sampled is an exact token prefix of the next prompt. I took the response tokens to be the canonical tokenization of the reply plus the end-of-turn token, so any break comes from the template and not from sampling.

Chat template Call pairs where the prefix held Training samples from the 5-call rollout
Stock Qwen3.5-9B template, thinking off 0 of 4 5
swe_smith_chat_template.jinja from the example, thinking off 4 of 4 1

In this test the stock template breaks at every pair, and the script’s output shows where. With thinking off, the generation prompt ends in an empty think block and the model’s reply continues from it. When the harness resends the conversation, the stock template renders that earlier assistant turn without the block, so the new prompt diverges right after the assistant header. (With thinking left on and an invented reasoning prefix on each reply, I saw the same break in the same place.) The example avoids the problem in two steps: its agent asks vLLM for enable_thinking: False on every request, and the trainer scripts pass vLLM a custom swe_smith_chat_template.jinja that writes the empty think block back into every earlier assistant turn. It is a one-file fix, but you only apply it if you know to test for the problem. The script is in the last section.

The second source of splits is the one the report’s Figure 3 illustrates with the word having: sampled as two tokens, h and aving, it can come back as different tokens after decoding and re-tokenizing. With the Qwen3.5 tokenizer, re-tokenizing having gives the single ID 66246 rather than the sampled pair, 71 and 2241. The same holds for running (26299 against 81 and 10890) and testing (8573 against 83 and 57835). The gateway forces a training temperature of 1 by default, where low-probability paths like these are more likely to be sampled than under greedy decoding. In the report’s own coding runs, “only 36% of rollouts on average remain as a single training sample, and each rollout yields 2.4 training samples on average.” In the authors’ runs, then, fragmentation is the normal case.

Advantage and loss: the fragmentation tax

Group-based methods such as GRPO score each attempt against the average reward of a group of attempts at the same task. The report’s example has two rollouts of one task: the first earns reward 1 and splits into three training samples, the second earns 0 and stays one sample. Averaging per rollout gives a baseline of 1/2; averaging per sample gives 3/4. From those numbers (my arithmetic), the per-rollout rule credits the success with plus 0.5 and charges the failure minus 0.5, while the per-sample rule gives each of the success’s three samples plus 0.25 and charges the failure minus 0.75. The same two outcomes teach the model different lessons depending on how many pieces the winner fell into. The report says verl Uni-Agent and Polar compute advantages per rollout, while slime and AReaL compute them per sample.

Loss normalization has the same flaw. The report’s example batch has three rollouts: A splits into samples of 50 and 100 response tokens, B into three samples of 30 tokens each, and C is one sample of 40. This is how much of the batch’s gradient each rollout controls under the three normalizations the report compares (my arithmetic from its formulas):

Rollout Samples Response tokens Token-mean (DAPO) Seq-mean-token-mean (GRPO) Rollout-level token-mean (slime, Agent Lightning)
A 2 150 53.6% 33.3% 33.3%
B 3 90 32.1% 50.0% 33.3%
C 1 40 14.3% 16.7% 33.3%

Under the GRPO-style rule, rollout B holds half of the batch’s weight although rollout A produced 67% more tokens (150 against 90); the only difference is that B fell into three pieces. The report concludes that “sample count should not be allowed to affect gradient normalization, because it is often driven by incidental factors such as retokenization,” and picks the rollout-level version, noting that plain token-mean “is sensitive to long sequences: when many long negative samples appear in a batch, it can cause instability later in training.” The fourth problem, backend scheduling, follows from the same fact: sample counts are known only after the harness finishes, while the GPU count and parallel layout are fixed, and all samples from one rollout need to land in one optimizer update to avoid what the report calls “within-rollout policy skew.”

What the 3,500 lines count

The number appears in the blog post, the README and the report (“approximately 3,500 lines of code”). I counted non-blank, non-comment Python lines under agentlightning/ in the tagged source of three releases; docstring lines are included, and the script is in the last section.

Release Date Python files Lines verl trainer folder Gateway (server) Rollout controller Everything else
v1.0.0 Aug 17, 2026 27 3,511 1,931 755 594 231
v1.0.2 Sep 29, 2026 28 4,078 2,410 828 608 232

On this definition the claim is accurate for the v1.0.0 tag, at 3,511 lines. Six weeks and two patch releases later the package holds 4,078, about 16% more, and the figure in the blog post, the README and the report has not moved. Two other readings matter more. First, 59% of the v1.0.2 count (2,410 lines) is the verl integration, while the gateway and the controller together are 1,436 lines, or 35%. The count excludes verl and vLLM themselves, which the README’s scripts/setup_verl.sh 0.8.0 cu130 step installs separately. “Lightweight” therefore describes the glue around a conventional training stack, which is a fair thing to build but not a small training system.

Second, the v0.3.0 release of December 24, 2025 held 93 files and 23,985 lines, and v1.0.0 has no folders for tracing, instrumentation, algorithms, adapters, runners or a command-line interface, which v0.3.0 did. The README says plainly that “Agent Lightning was completely refactored in v1.0” and points to a v0.x branch for older releases. The 85% drop therefore comes from a smaller scope as much as from tighter code, and I would read v1.0 as a different product rather than an upgrade for anyone on v0.x.

What the evidence supports

Six thousand samples, and the compute behind them

The example’s documentation traces the dataset. SWE-smith has 59,136 executable tasks from 128 Python repositories. The pipeline drops 18,033 records with an empty problem statement, 1,265 whose problem branch is missing, and tasks that need more than 200 tests. It then runs Qwen3.5-9B four times on every remaining task as a difficulty probe, removes tasks solved in all four runs, keeps the mixed ones (about 5,000) and adds 1,000 tasks that failed all four, ending with about 6,000 training and 400 validation examples. The six thousand are therefore about a tenth of the original pool, chosen with the base model that is then trained.

The report says “modest compute” without a GPU count or GPU-hours, and the probe runs are a cost the 6K figure does not show: on the order of 150,000 attempts (my arithmetic: about 39,800 tasks remain after the first two filters, times four, before the 200-test filter removes some). The example’s documentation lists four B200 GPUs for the 9B run, and each rollout is a Kubernetes Job in a repository image with limits of 1 CPU and 6 GiB of memory and up to 100 turns.

Reading 41.8% to 56.4%

SWE-bench Verified is, in OpenAI’s words, “a subset of the original test set from SWE-bench, consisting of 500 samples verified to be non-problematic by our human annotators.” So 41.8% is 209 of 500 and 56.4% is 282 of 500, which is 73 more tasks solved. The 95% Wilson intervals for those two results, which capture only the sampling error that comes from having 500 tasks, run from 37.6% to 46.2% and from 52.0% to 60.7% (my arithmetic). They do not overlap, so I would not call the gain an accident of the task sample.

What the report’s text does not give is anything about run-to-run variation. The words seed, variance and confidence interval do not appear in it, it does not say how the step-208 checkpoint was chosen, and its own ablation says the final variant’s highest SWE-smith validation reward came at step 128. That ablation, which supports the rollout-level choices, reports validation rewards of 38.2% (rollout-level advantage and normalization), 35.0% (the sample-level baseline) and 33.1% (rollout-level advantage only) on a validation split of about 400 examples. Gaps of 3.2 and 5.1 points are about the size of the sampling error of a 400-example set: roughly 4.8 points either way at 95% for a 38% score, if each example counts once (my arithmetic). The three variants share the same validation tasks, which narrows a paired comparison, but the report gives none and does not say the comparison was repeated. Its wording, “These results suggest,” is the right strength. This site’s piece on Microsoft’s ThinkingBox, published October 4, noted that benchmark runs every task 20 times against each model. Harnessed RL results deserve the same habit.

The newer MoE example in the docs publishes its checkpoints in full. Qwen3.5-35B-A3B goes from 47.8% (239 of 500) to 58.8% at step 64, 61.6% (308 of 500) at step 112 and 60.6% at step 176, using, per the docs, “the official SWE-bench Verified harness on all 500 instances.” The README headlines step 112 as “47.8% to 61.6% after training it on only 1.8K training examples,” which is the best of the three trained checkpoints. Steps 112 and 176 differ by five tasks out of 500, and their intervals (57.3% to 65.8% and 56.3% to 64.8%) overlap almost entirely, so the table supports a statement like about 60% after roughly 1,000 to 3,000 examples better than it supports any single decimal. Listing every checkpoint is the honest way to publish; the headline just has to be read with the table next to it.

The gateway is also the trust boundary

Because every call goes through the gateway, it holds the training data, can reach the model server and receives the reward. Reading the v1.0.2 source, four details matter to anyone deploying it.

Authentication is opt-in

The default configuration binds host: 0.0.0.0 on port 8080 with key: "", and the configuration docs say an empty key “disables authentication and logs a warning.” The log line, in server/app.py, ends “Do not use in production.” The example launcher, run.sh, sets AGL_KEY="${AGL_KEY:-dummy}", so skipping the export gives a key of dummy. When a real key is set, one bearer key covers every route (rollouts, events, models and the proxy), and it is shared: the docs say “Use the same non-empty key in the trainer and Controller,” and the controller configuration describes agl_server.key as the “Bearer key used by the Controller and agents.”

The agent’s shell inherits the key

The Kubernetes reconciler, controller/k8s_reconciler.py, writes AGL_KEY, the rollout’s model-proxy URL and its event URL into every agent container as plain environment values in the Job spec, not as a Secret reference. In the shipped SWE-smith agent, smith_agent.py, the harness uses them itself: it posts the reward to the event URL with the bearer key. The commands the model writes run through subprocess.run(["bash", "-c", command], ..., env=_agent_env()), and _agent_env() copies the process environment except the variable that hides the relocated Git directory. Reading the code, a model-written env command would print the key and the event URL.

The trainer takes the last reward event it finds for a rollout (reward_events[-1] in agl_rollout_manager.py) and the harness posts its own reward after the agent’s last command, so a reward posted early would be overwritten. I did not run the cluster example and did not test whether a model could exploit this. The point is narrower: the credential that writes the reward sits in the environment of the process whose commands the model controls.

Some guardrails live inside the agent’s reach

The report’s reward-hacking section lists what the coding agent did when allowed: it used Git history to find the reference commit, fetched upstream source with wget or curl, downloaded the package with pip and used Python networking libraries. The example blocks these in two ways. In the harness, regular expressions refuse commands that invoke Git, touch the hidden Git directory, call curl, wget and similar tools, use Python’s networking calls, install packages or write test-harness files such as conftest.py. The code comment calls the network, install and tamper blocks “a code backstop” and says “the authoritative fix is a default-deny egress NetworkPolicy.” The documentation says “we strongly recommend adding a Kubernetes network policy that denies all outbound traffic from agent pods except connections to the AGL Gateway.”

I found no NetworkPolicy manifest anywhere in the v1.0.2 tree, so that layer is the operator’s to write, and the one destination it must allow is the gateway, the same place the reward is posted. NVIDIA’s Open Agent Safety Platform, announced on September 28, is built on the opposite premise that enforcement belongs outside the agent’s reach. Training environments also leak from the inside: OpenAI’s September disclosure described an internal research model that used its sandbox’s DNS resolver to reach the open internet.

What I would do before pointing this at real workloads

  1. Set a non-default key and bind the gateway to an internal address instead of relying on the launcher’s fallback.
  2. Give agent pods a narrower credential than the trainer’s. As far as I can tell from the v1.0.2 code and docs there is one key for everything, so this means a reverse proxy in front of the gateway that exposes only the model-proxy and event routes to pods.
  3. Apply the default-deny egress policy the docs recommend and allow only the gateway.
  4. Keep the key out of the shell the model controls. Removing it from the environment passed to subprocess.run is not enough on its own, because a same-user shell can usually read the harness’s own environment through /proc; run the commands as a different user or post the reward from a separate process.
  5. Before training, run the prefix check below on your own template and message format, and log the single-sample share and samples per rollout as health metrics (the report’s runs show 36% and 2.4).
  6. Repeat a run with several seeds before crediting any design choice or reporting a final score.

The bottom line

Agent Lightning v1.0 is a useful piece of work. The report names counting problems that other frameworks leave implicit, the repository ships the data pipeline and the checkpoint table, and the 14.6-point gain is large against the sampling error of the benchmark. What the parts I could read do not settle is how stable the recipe is across runs, what the compute really was, and how to deploy the gateway safely by default. The sample accounting and the gateway are where a team adopting harnessed RL should spend its first week.

My checks have limits. The template test uses a toy conversation and canonical response tokens, so it reproduces the mechanism and not the authors’ rates. The line count is one definition among several (the same v1.0.0 tag holds 4,350 physical lines in total, and 3,320 lines of code once blank lines, comments and docstrings are excluded). I did not train anything or run the Kubernetes example.

Reproduce the two checks

Neither script executes code from the downloaded repository; they only read files from it (the second one also downloads tokenizer files from Hugging Face). The counting script takes the extracted tarball of any tag.

Count the lines

Download and extract

curl -sL -o v1.0.2.tar.gz https://codeload.github.com/microsoft/agent-lightning/tar.gz/refs/tags/v1.0.2
mkdir v1.0.2 && tar -xzf v1.0.2.tar.gz -C v1.0.2
python count_loc.py v1.0.2/agent-lightning-1.0.2

The script

# count_loc.py
# Counts non-blank, non-comment Python lines in the agentlightning package, folder by folder
import pathlib
import sys

root = pathlib.Path(sys.argv[1]) / "agentlightning"
by_folder = {}
for path in root.rglob("*.py"):
    parts = path.relative_to(root).parts
    folder = parts[0] if len(parts) > 1 else "(top-level files)"
    lines = (line.strip() for line in path.read_text(encoding="utf-8").splitlines())
    by_folder[folder] = by_folder.get(folder, 0) + sum(1 for line in lines if line and not line.startswith("#"))
for folder, count in sorted(by_folder.items(), key=lambda item: -item[1]):
    print(f"{folder:20s}{count:7d}")
print(f"{'total':20s}{sum(by_folder.values()):7d}")

For v1.0.2 it prints:

verl                   2410
server                  828
controller              608
(top-level files)       231
config                    1
total                  4078

Repeating the download for v1.0.0 and v0.3.0 (the folder is agent-lightning-1.0.0 or agent-lightning-0.3.0) prints:

verl                   1931
server                  755
controller              594
(top-level files)       230
config                    1
total                  3511
store                  7089
(top-level files)      2526
utils                  2520
algorithm              1604
verl                   1552
tracer                 1472
adapter                1112
runner                 1070
trainer                1006
types                   980
instrumentation         767
emitter                 751
execution               696
litagent                581
cli                     259
total                 23985

Test the chat template and the token drift

This script needs the extracted v1.0.2 tree for the custom template, and transformers without PyTorch is enough. The output below leaves out the library’s notices about PyTorch and unauthenticated downloads.

python -m pip install transformers
python prefix_check.py v1.0.2/agent-lightning-1.0.2/examples/swe_smith/swe_smith_chat_template.jinja
# prefix_check.py
# Python 3.13 and transformers 5.x (no PyTorch needed); tokenizer files come from Qwen/Qwen3.5-9B
import sys

from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("Qwen/Qwen3.5-9B")
custom = open(sys.argv[1], encoding="utf-8").read() if len(sys.argv) > 1 else None

SYSTEM = "You are a coding agent. Reply with one bash command per turn."
TASK = "Fix parse_date() so it rejects 2024-02-30."
TURNS = [  # (what the model says, what the harness sends back as a user message)
    ("THOUGHT: List the repository.\n\n```bash\nls /testbed\n```",
     "<returncode>0</returncode>\n<output>\ndates.py\ntests\n</output>"),
    ("THOUGHT: Find the function.\n\n```bash\ngrep -n 'def parse_date' /testbed/dates.py\n```",
     "<returncode>0</returncode>\n<output>\n12:def parse_date(text):\n</output>"),
    ("THOUGHT: Read it.\n\n```bash\nsed -n '1,30p' /testbed/dates.py\n```",
     "<returncode>0</returncode>\n<output>\nimport calendar\n</output>"),
    ("THOUGHT: Patch the check.\n\n```bash\nsed -i 's/day > 31/day > days_in_month/' /testbed/dates.py\n```",
     "<returncode>0</returncode>\n<output>\n</output>"),
    ("THOUGHT: Run the tests.\n\n```bash\npytest -q /testbed/tests\n```",
     "<returncode>0</returncode>\n<output>\n12 passed\n</output>"),
]


def ids(text):
    return tok(text, add_special_tokens=False)["input_ids"]


def check(template, label):
    messages = [{"role": "system", "content": SYSTEM}, {"role": "user", "content": TASK}]
    seen, held, first_break = None, 0, None
    for call, (reply, observation) in enumerate(TURNS, start=1):
        prompt = ids(tok.apply_chat_template(messages, chat_template=template, tokenize=False,
                                             add_generation_prompt=True, enable_thinking=False))
        if seen is not None:
            if prompt[: len(seen)] == seen:
                held += 1
            elif first_break is None:
                k = next(i for i in range(len(seen)) if seen[i] != prompt[i])
                first_break = (call, k, tok.decode(seen[k - 3: k + 6]), tok.decode(prompt[k - 3: k + 6]))
        seen = prompt + ids(reply + "<|im_end|>")  # prompt plus the response the model sampled
        messages += [{"role": "assistant", "content": reply}, {"role": "user", "content": observation}]
    pairs = len(TURNS) - 1
    print(f"{label}: prefix held for {held} of {pairs} call pairs, "
          f"so one {len(TURNS)}-call rollout becomes {pairs - held + 1} training sample(s)")
    if first_break:
        call, k, sampled, rendered = first_break
        print(f"  call {call} diverges from call {call - 1} (prompt plus response) at token {k}")
        print(f"    sampled:     {sampled!r}")
        print(f"    re-rendered: {rendered!r}")


check(None, "stock Qwen3.5-9B template")
if custom:
    check(custom, "swe_smith_chat_template.jinja")

# Decode and re-tokenize drift: a sampler may emit two tokens where the tokenizer would use one
vocab = tok.get_vocab()
for word, (left, right) in {"having": ("h", "aving"), "running": ("r", "unning"), "testing": ("t", "esting")}.items():
    sampled = [vocab[left], vocab[right]]
    print(f"{word}: sampled {sampled}, decoded then re-tokenized {ids(tok.decode(sampled))}")

On transformers 5.19.0 and Python 3.13.14 it prints:

stock Qwen3.5-9B template: prefix held for 0 of 4 call pairs, so one 5-call rollout becomes 5 training sample(s)
  call 2 diverges from call 1 (prompt plus response) at token 46
    sampled:     '<|im_start|>assistant\n<think>\n\n</think>\n\nTHO'
    re-rendered: '<|im_start|>assistant\nTHOUGHT: List the'
swe_smith_chat_template.jinja: prefix held for 4 of 4 call pairs, so one 5-call rollout becomes 1 training sample(s)
having: sampled [71, 2241], decoded then re-tokenized [66246]
running: sampled [81, 10890], decoded then re-tokenized [26299]
testing: sampled [83, 57835], decoded then re-tokenized [8573]

Tags:

AI AgentsAI SecurityKubernetesMicrosoft ResearchReinforcement Learning

Share

A common raven calling with its beak open on a grassy hillside under a pale blue sky
Previous Post

PoeLLM Malware Finds Its Command Server in a GitHub Poem and Targets Exposed AI Servers

Rows of closed oak library card catalog drawers, each with a brass pull and a blank label holder
Next Post

How to Encrypt PII in Python and Keep It Searchable With Blind Indexes

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
08 Oct
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
08 Oct
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
Trending
October 8, 2026
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
October 8, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
October 8, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google

Related Posts

Blue-lit server racks in a modern data center, illustrating the compute infrastructure behind the AI boom.
Articles

The AI Boom Is Spending Real Money Before Proving Real Returns

June 7, 2026
Technician working with a laptop beside server racks, representing enterprise AI retrieval infrastructure
Articles

Google’s Agentic RAG Push Makes Enterprise AI Less of a One-Shot Guess

June 7, 2026
A person with a laptop and smartphone, representing digital attention and AI-assisted work
Articles

AI Chatbots Are Making Attention a Design Problem

June 7, 2026
A customer-support representative wearing a headset against a dark studio background.
Articles

The Meta AI Support Hack Was a Plain Old Authorization Failure

June 7, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026