TRENDING
Rows of identical brass-colored apartment mailboxes with small locks and name labels along an orange corridor wall
October 9, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
Street-level upward view of the Monetary Authority of Singapore building and neighbouring office towers under a pale sky
October 9, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
Cast-iron late Qing dynasty coin minting press with a large flywheel, displayed in a museum case
October 9, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google
Rows of closed oak library card catalog drawers, each with a brass pull and a blank label holder
October 9, 2026
How to Encrypt PII in Python and Keep It Searchable With Blind Indexes
Close-up of a vintage Western Electric manual telephone switchboard with orange lamps, red patch cords plugged into jacks, a rotary dial and a black handset
October 9, 2026
Microsoft’s Agent Lightning v1.0 Turns Agent Training Into a Sample-Accounting Problem
09 Oct 2026
SXZ.io SXZ.io
  • Home
Search the Site
Popular Searches:
Technology Amazon AI
Recent Posts
Two orange safety relief valves on grey pressure vessels in an industrial plant
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
Yellow diamond-shaped merging traffic warning sign showing a side road joining a main road
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A lugworm lying on wet sand and mud at low tide
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
SXZ.io SXZ.io
  • Home

Categories

Articles 232 Posts
News 234 Posts
Learning Hub 204 Posts
Home/News/NVIDIA’s Nemotron 3.5 Lightning Ships With Day-One Ubuntu Support From Canonical
News

NVIDIA’s Nemotron 3.5 Lightning Ships With Day-One Ubuntu Support From Canonical

NVIDIA opened its Nemotron 3.5 Lightning agent model and NeMo Switchyard router on August 11, and Canonical shipped the model on Ubuntu the same day through a single-command inference snap.

August 11, 2026 4 Min Read
38

NVIDIA released Nemotron 3.5 Lightning on August 11, an open 30-billion-parameter mixture-of-experts model built for long-running AI agents, alongside NeMo Switchyard, an open source library for routing agent requests between Lightning and larger frontier models. Canonical made the model available on Ubuntu the same day, packaging it as a single-command inference snap.

Table Of Content

  • A Small Model Built to Run Agents All Day
  • NeMo Switchyard Routes Work Down From Frontier Models
  • Canonical Ships It on Ubuntu the Same Day

A Small Model Built to Run Agents All Day

Nemotron 3.5 Lightning has 30 billion total parameters but only 3 billion active at a time, using a hybrid architecture that interleaves Mamba-2 layers, Mixture-of-Experts layers, and a smaller number of attention layers. NVIDIA’s technical blog post describes it as the smallest member of the Nemotron 3 family, built as a fast, specialized worker for narrow tasks inside larger multi-agent systems rather than a general-purpose chat model. It supports a context window of up to 1 million tokens, letting an agent hold a long, multi-step task in memory without losing earlier steps.

NVIDIA says the model delivers up to 4 times the output speed of similarly sized models and finishes agentic tasks 30% faster than competitors in its class, according to the company’s announcement. On PinchBench, an agentic benchmark cited in NVIDIA’s technical post, Lightning reached 86% accuracy while finishing 10,000 tasks 30% faster than Qwen3.6 35B at a similar accuracy level. On more conventional benchmarks listed on the model’s Hugging Face model card, it scores 81.62 on MMLU Pro, 75.57 on GPQA Diamond, 52.80 on SWE-bench Verified, and 72.88 on IFBench under loose grading.

Independent measurement backs up some of that. Artificial Analysis scores the model at 24 on its Intelligence Index, a 9-point jump over the smaller Nemotron 3 Nano’s score of 15 and roughly in line with OpenAI’s gpt-oss-120b. The same analysis puts Lightning’s GDPval-AA v2 Elo at 824, ahead of both Nemotron 3 Super and gpt-oss-120b, and measured its Terminal-Bench v2.1 score at 24%, well above Nemotron 3 Nano’s 7%. Artificial Analysis also clocked close to 670 tokens per second on a DeepInfra endpoint and roughly 0.5 minutes per Intelligence Index task, versus about 3.5 minutes for the similarly sized Qwen3.6 35B.

NVIDIA released the model under its OpenMDW-1.1 license, which the company says allows commercial use, modification, and redistribution without asking NVIDIA’s permission. It is available now on Hugging Face, ModelScope, and OpenRouter, and as a hosted NIM microservice through build.nvidia.com. NVIDIA lists support across vLLM, SGLang, TensorRT-LLM, Ollama, llama.cpp, LM Studio, and Unsloth, with hardware targets ranging from a single RTX PC or DGX Spark up through data center GPUs; NVFP4-quantized and speculative-decoding checkpoints ship alongside the full-precision BF16 release for teams that want to trade some accuracy for speed.

NeMo Switchyard Routes Work Down From Frontier Models

Alongside the model, NVIDIA released NeMo Switchyard, an open source routing library that sits in front of an agent and decides which model should handle each step of a task based on the quality, latency, and cost that step requires. NVIDIA frames the pattern as sending planning steps up to a frontier model and pushing routine execution steps down to Lightning, and says the combination “maintains frontier-level accuracy while reducing task completion cost to nearly one-third of Opus 4.8 alone,” a reference to Anthropic’s Opus 4.8 model. Switchyard’s code is on GitHub now, and NVIDIA says support is coming to partner platforms including Boomi, Cadence, Classmethod, Cognition, Kong, LangChain, LiteLLM, Nous Research, Ramp, and Siemens.

NVIDIA says several companies have already customized Lightning for production use, including CrowdStrike in cybersecurity, Harvey working with Trajectory in legal work, Lila Sciences in physical and life sciences, and Fastino Labs across finance, healthcare, and software development. In comments reported by SiliconANGLE, NVIDIA vice president of generative AI Kari Briski said Lightning is “remarkably easy to customize.” She pointed to CodeRabbit, which used NVIDIA’s standard auto model recipe to train a router agent for $85 in around two hours, and described other unnamed partners training on a single H100 GPU relatively inexpensively, dropping Lightning into an existing post-training stack “with no changes required,” and setting up an overnight training job and returning the next morning to collect the results.

Canonical Ships It on Ubuntu the Same Day

Canonical’s own announcement, also published August 11, says Nemotron 3.5 Lightning is available on Ubuntu at launch through inference snaps: pre-packaged AI inference runtimes distributed as snap packages for consistent deployment across systems. Installing the model takes one command, sudo snap install nemotron-3-5-lightning, and Canonical says the resulting deployment behaves the same way on workstations, edge devices, and servers. The pitch to enterprises is standardization: snap packaging handles confinement, verified distribution, and automatic updates, so teams deploying the model do not have to build and maintain their own inference infrastructure around it.

Pairing a small, fast, cheaply customized model with an installable, self-updating runtime targets the operational side of running agents in production, where the cost of maintaining inference infrastructure across a fleet of machines can outweigh the cost of the model itself.

Tags:

AI AgentsCanonicalnemotronNVIDIAUbuntu

Share

Archery target with a tight cluster of arrows landing in the bullseye, illustrating retrieval precision
Previous Post

How to Evaluate RAG Retrieval Quality With Precision, Recall, and MRR in Python

Floodwater rushes through an open spillway gate at the Narayanpur Dam on the Krishna River in Karnataka, India
Next Post

Cloudflare’s H1 2026 DDoS Report Turns the News Cycle Into an Attack Signal

No Comment! Be the first one.

Leave a Reply Cancel reply

Your email address will not be published. Required fields are marked *

Latest
08 Oct
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
08 Oct
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
Trending
October 8, 2026
How to Add Backpressure and Load Shedding to a Python Service Before Overload Takes It Down
October 8, 2026
GitHub’s Git Rebuild Turns Repository Durability and Read Scale Into Two Separate Problems
October 8, 2026
A Compromised Admin Account Put the Shai-Hulud Worm Into AI Sandbox Maker Tensorlake’s npm SDK
October 8, 2026
How to Prevent Broken Object Level Authorization (IDOR) in a FastAPI App
October 8, 2026
Singapore’s AI Guidelines Turn Independent Review Into a Question of Who Sets the Risk Rating
October 8, 2026
Attackers Hijacked the .gh, .sl and .as Country Domains and Minted HTTPS Certificates for Google

Related Posts

Rows of server racks in a data center representing network infrastructure targeted by botnets
News

C0XMO Botnet Shows Why Old Router Firmware Still Matters

June 7, 2026
Close-up of a USB flash drive, representing physical data-theft risk in office security incidents
News

Fake IT Support Is Now Walking Through the Front Door

June 7, 2026
A phone security app on a smartphone resting on a laptop keyboard.
News

Everest Forms Pro Flaw Is Being Exploited to Create Rogue WordPress Admins

June 7, 2026
Rows of server racks in a data center, illustrating the infrastructure behind frontier AI funding.
Articles

AI’s Biggest Backers Are Hedging the Frontier Model Race

June 8, 2026
SXZ.io SXZ.io
  • [email protected]

Categories

Articles
Learning Hub
News

All Rights Reserved by SXZ.io ©2026