ModelsAgents 🇺🇸 11.08.2026 16:01

NVIDIA Introduces Nemotron 3.5 Lightning for Fast, Specialized Agentic Tasks

NVIDIANVIDIA
NVIDIA has expanded its Nemotron 3 model family with Nemotron 3.5 Lightning, an open 30B mixture-of-experts model designed for always-on agents, offering up to 4x faster token generation and 30% faster time to completion compared to open models in its class. The model can be fine-tuned for personalized tasks and runs locally on various NVIDIA platforms. NVIDIA also introduced NeMo Switchyard, an open-source routing library that optimizes agent workflows for accuracy, speed, and cost, reportedly reducing benchmark completion costs to roughly one-third of Opus 4.8 alone.
NVIDIA has announced Nemotron 3.5 Lightning, a customizable open 30B mixture-of-experts (MoE) model for always-on agents, expanding its Nemotron 3 family. The model delivers up to 4x faster token generation and 30% faster time to completion compared to open models in its class. Because it has open weights, developers can fine-tune it for specific styles, specialties, or coding conventions, enabling personalized local agentic AI experiences such as email management, smart-home routines, or coding companions. NVIDIA collaborated with vLLM, Ollama, llama.cpp, LM Studio, and Unsloth to provide optimized local deployment options, including NVFP4 and GGUF formats. The model runs locally on devices ranging from NVIDIA RTX PCs and DGX Spark to Jetson, and scales to workstations and data centers with Blackwell systems from partners like Acer, ASUS, Dell, and Lenovo. Additionally, NVIDIA introduced NeMo Switchyard, an open-source routing library that directs each step of an agent workflow to the best-fit model based on accuracy, speed, and cost, helping enterprises manage token costs. Internal benchmarks show NeMo Switchyard reduced benchmark completion costs to roughly one-third of Opus 4.8 while maintaining frontier-level task completion. NeMo Switchyard is available on GitHub.
Abbreviations
MoE = Mixture of Experts — Микс экспертов
RTX = Ray Tracing Texel eXtreme — трассировка лучей
Source: NVIDIA blog — original
Our earlier posts on this topic ↓
Fresh news