Thursday, August 13, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

Nvidia announced new Nemotron models and NeMo router software, advancing open-source AI deployment and reducing dependency on closed ecosystems.

Nvidia's integrated model-plus-router strategy enables customers to optimize end-to-end inference pipelines, capturing more infrastructure value.
Trade pressSlicast · August 12, 2026 · Global · Source: The Next Platform
importance 66

Nvidia continues to expand its influence across the AI ecosystem through technological expansion and strategic partnerships. The vendor is adding to its portfolio of open AI models and launching an open-source model library that helps organizations select the right model for each step in an agentic AI workflow—eliminating the need to default to a single model or manually evaluate options.

In the model space, Nvidia is expanding its Nemotron family with Nemotron 3.5 Lightning, designed to operate as part of what the company calls "systems of models," or model ensembles where different models handle different workflow steps. This represents a layer of abstraction beyond mixture-of-experts (MoE) approaches. Kari Briski, vice president of generative AI at Nvidia, explained: "Efficient agents need a system of models, not just one model. Agent solutions and workflows make many model calls. After a plan is devised, it is carried out in steps, or what we call turns. Different steps require many tasks at different levels of intelligence."

Nvidia has rolled out three prior Nemotron models over the past year. Nemotron Nano, released in December 2025, is a high-throughput, efficient small language model designed for real-time AI agents, multi-step tool calls, advanced math, and coding. Nemotron 3 Super is a 120-billion parameter open-weight model with an MoE design for efficient agentic workflows, complex reasoning, and accurate enterprise tool calling. Nemotron 3 Ultra, a reasoning model with 550-billion parameters, handles high-throughput and long-running agents for multi-step workflows including agent orchestration, deep repository coding, research, and enterprise automation.

Lightning combines capabilities across this family. "Lightning brings together the best of our family: the knowledge of Ultra packed into the size of our Nano," Briski said. "Lightning is built for always-on agents handling a constant stream of specialized tasks—a supercharged workhorse model for high-volume AI that is fast, accurate, and efficient, delivering four times the throughput in its class at the same intelligence index as our Super model." The model can run on Nvidia's Jetson, GeForce RTX, DGX Spark, and DGX Station systems, allowing developers to run operations on their own hardware rather than calling cloud APIs for every task.

According to Nvidia's testing via PinchBench—which rates open models on agentic tasks like coding, research, and file management—Lightning is up to 30 percent faster than comparable Qwen models and faster than models like Google DeepMind's Gemma. The 4X output speed results in a 30 percent speedup in completing agentic AI tasks. Organizations can post-train Lightning via the NeMo platform on proprietary data and tools. Nvidia is publishing its training data and techniques, and releasing Nemotron-RL-Agentic-Terminal-Pivot, an open dataset for training coding agents.

Lightning leverages Nvidia's hybrid Mamba transformer architecture and multi-token prediction technology. The model has outperformed other open and proprietary models in various tasks involving CodeRabbit, Harvey, and Lila, and performed well against Nemotron 3 Super for CrowdStrike.

Alongside Lightning, Nvidia released the NeMo Switchyard routing library, enabling organizations to select the optimal model for each task in an agentic workflow. Briski noted: "The best model changes as the workflow evolves. An agent has many states when completing tasks. The agent state changes as the tools return results, errors occur, or some tasks become routine. The router continuously balances quality, latency, and cost. A fixed model choice can't adapt to those changes." AI model routers aren't new—vendors including OpenRouter, Not Diamond, and LiteLLM offer similar capabilities. Microsoft unveiled Switchcraft in May, claiming it increases accuracy by nearly 83 percent and reduces inference costs by 84 percent.

Nvidia positions Switchyard alongside its open model lineup to provide developers with integrated workflow and routing capabilities. "Switchyard gives developers and platforms a way to define their own model pools, routing criteria, and policies," Briski explained. "It routes frontier models to reasoning-intensive steps and Lightning to execution tasks where speed and efficiency matter. That is the power of a system of models—matching the right model to each step of the workflow." Testing shows Switchyard delivers better accuracy at one-third the cost of Anthropic's Opus 4.8 alone, with improvements when combining Opus 4.8 with Lightning and Gemma 3.

Beyond models, Nvidia is advancing AI security. Following reports of autonomous breaches of Hugging Face involving OpenAI's GPT-5.6-Sol and other unreleased frontier models, Nvidia led development of the Open Secure AI Alliance—a consortium of over three dozen companies arguing that protections against AI threats must be built on open models and tools all defenders can adapt and deploy. The consortium now includes over 120 members and is developing guidelines for enhanced agentic AI cybersecurity. Nvidia is also staffing an AI safety and security team.

To address AI infrastructure costs, Nvidia is partnering with Apollo, BlackRock, Blackstone, Brookfield, Goldman Sachs, and KKR to establish a $500 billion fund supporting enterprise, startup, and government AI infrastructure investment. Jensen Huang, Nvidia co-founder and chief executive officer, stated: "In AI, compute is revenue. Our DSX AI factories and broad ecosystem of developers, customers and offtakers demonstrate why we are bringing together the world's leading long-term capital providers to independently underwrite AI infrastructure."

Read the original
Nvidia announced new Nemotron models and NeMo… · Slicast