NVIDIA’s Nemotron 3.5 Lightning Ships With Day-One Ubuntu Support From Canonical
NVIDIA opened its Nemotron 3.5 Lightning agent model and NeMo Switchyard router on August 11, and Canonical shipped the model on Ubuntu the same day through a single-command inference snap.
NVIDIA released Nemotron 3.5 Lightning on August 11, an open 30-billion-parameter mixture-of-experts model built for long-running AI agents, alongside NeMo Switchyard, an open source library for routing agent requests between Lightning and larger frontier models. Canonical made the model available on Ubuntu the same day, packaging it as a single-command inference snap.
Table Of Content
A Small Model Built to Run Agents All Day
Nemotron 3.5 Lightning has 30 billion total parameters but only 3 billion active at a time, using a hybrid architecture that interleaves Mamba-2 layers, Mixture-of-Experts layers, and a smaller number of attention layers. NVIDIA’s technical blog post describes it as the smallest member of the Nemotron 3 family, built as a fast, specialized worker for narrow tasks inside larger multi-agent systems rather than a general-purpose chat model. It supports a context window of up to 1 million tokens, letting an agent hold a long, multi-step task in memory without losing earlier steps.
NVIDIA says the model delivers up to 4 times the output speed of similarly sized models and finishes agentic tasks 30% faster than competitors in its class, according to the company’s announcement. On PinchBench, an agentic benchmark cited in NVIDIA’s technical post, Lightning reached 86% accuracy while finishing 10,000 tasks 30% faster than Qwen3.6 35B at a similar accuracy level. On more conventional benchmarks listed on the model’s Hugging Face model card, it scores 81.62 on MMLU Pro, 75.57 on GPQA Diamond, 52.80 on SWE-bench Verified, and 72.88 on IFBench under loose grading.
Independent measurement backs up some of that. Artificial Analysis scores the model at 24 on its Intelligence Index, a 9-point jump over the smaller Nemotron 3 Nano’s score of 15 and roughly in line with OpenAI’s gpt-oss-120b. The same analysis puts Lightning’s GDPval-AA v2 Elo at 824, ahead of both Nemotron 3 Super and gpt-oss-120b, and measured its Terminal-Bench v2.1 score at 24%, well above Nemotron 3 Nano’s 7%. Artificial Analysis also clocked close to 670 tokens per second on a DeepInfra endpoint and roughly 0.5 minutes per Intelligence Index task, versus about 3.5 minutes for the similarly sized Qwen3.6 35B.
NVIDIA released the model under its OpenMDW-1.1 license, which the company says allows commercial use, modification, and redistribution without asking NVIDIA’s permission. It is available now on Hugging Face, ModelScope, and OpenRouter, and as a hosted NIM microservice through build.nvidia.com. NVIDIA lists support across vLLM, SGLang, TensorRT-LLM, Ollama, llama.cpp, LM Studio, and Unsloth, with hardware targets ranging from a single RTX PC or DGX Spark up through data center GPUs; NVFP4-quantized and speculative-decoding checkpoints ship alongside the full-precision BF16 release for teams that want to trade some accuracy for speed.
NeMo Switchyard Routes Work Down From Frontier Models
Alongside the model, NVIDIA released NeMo Switchyard, an open source routing library that sits in front of an agent and decides which model should handle each step of a task based on the quality, latency, and cost that step requires. NVIDIA frames the pattern as sending planning steps up to a frontier model and pushing routine execution steps down to Lightning, and says the combination “maintains frontier-level accuracy while reducing task completion cost to nearly one-third of Opus 4.8 alone,” a reference to Anthropic’s Opus 4.8 model. Switchyard’s code is on GitHub now, and NVIDIA says support is coming to partner platforms including Boomi, Cadence, Classmethod, Cognition, Kong, LangChain, LiteLLM, Nous Research, Ramp, and Siemens.
NVIDIA says several companies have already customized Lightning for production use, including CrowdStrike in cybersecurity, Harvey working with Trajectory in legal work, Lila Sciences in physical and life sciences, and Fastino Labs across finance, healthcare, and software development. In comments reported by SiliconANGLE, NVIDIA vice president of generative AI Kari Briski said Lightning is “remarkably easy to customize.” She pointed to CodeRabbit, which used NVIDIA’s standard auto model recipe to train a router agent for $85 in around two hours, and described other unnamed partners training on a single H100 GPU relatively inexpensively, dropping Lightning into an existing post-training stack “with no changes required,” and setting up an overnight training job and returning the next morning to collect the results.
Canonical Ships It on Ubuntu the Same Day
Canonical’s own announcement, also published August 11, says Nemotron 3.5 Lightning is available on Ubuntu at launch through inference snaps: pre-packaged AI inference runtimes distributed as snap packages for consistent deployment across systems. Installing the model takes one command, sudo snap install nemotron-3-5-lightning, and Canonical says the resulting deployment behaves the same way on workstations, edge devices, and servers. The pitch to enterprises is standardization: snap packaging handles confinement, verified distribution, and automatic updates, so teams deploying the model do not have to build and maintain their own inference infrastructure around it.
Pairing a small, fast, cheaply customized model with an installable, self-updating runtime targets the operational side of running agents in production, where the cost of maintaining inference infrastructure across a fleet of machines can outweigh the cost of the model itself.








No Comment! Be the first one.