ModelsAgents 🇺🇸 11.08.2026 16:01

NVIDIA launches Nemotron 3.5 Lightning and NeMo Switchyard for efficient agentic AI

NVIDIANVIDIA OpenAIOpenAI AnthropicAnthropic
NVIDIA introduced Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model optimized for high-volume agentic tasks, delivering up to 4x faster output and 30% faster task completion. The company also released NeMo Switchyard, an open-source model routing library for intelligent request distribution in agent workflows. Together they offer greater control over deployment and efficiency across various environments.
NVIDIA is expanding its Nemotron 3 model family with the release of Nemotron 3.5 Lightning, a 30-billion-parameter mixture-of-experts model designed for specialized tasks within larger multi-agent systems. The model delivers up to 4x faster output speed and 30% faster agentic task completion compared to others in its class. It is fully customizable and can be post-trained with NVIDIA NeMo on an organization's domain data, tools, and workflows. The company also released NeMo Switchyard, an open-source library for model routing that automatically directs prompts to the most capable and efficient model for each step of an agent workflow. Internal benchmarks show NeMo Switchyard maintains frontier-level accuracy while reducing task completion cost to nearly a third of Opus 4.8 alone. Partners including CrowdStrike, Harvey, CodeRabbit, Lila Sciences, and Fastino Labs are customizing Nemotron 3.5 Lightning for their workloads, while Boomi, Cadence, Cognition, Kong, LangChain, LiteLLM, Nous Research, Ramp, Siemens, and Classmethod are integrating or evaluating NeMo Switchyard.
Abbreviations
NIM = NVIDIA Inference Microservice — NVIDIA Inference Microservice
Source: NVIDIA blog — original
Our earlier posts on this topic ↓
Fresh news