Friday, September 11, 2026
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

NVIDIA is launching new hardware, software tools, and performance optimizations to make local AI inference faster and ea

NVIDIA official — first-hand confirmation of roadmap / product.
Official disclosureSlicast · September 11, 2026 · US · Source: NVIDIA Blog

NVIDIA, Microsoft and partners are collaborating to accelerate local AI inference and simplify agent deployment on consumer and professional NVIDIA hardware. The announcements include new compact RTX Spark Windows PCs from Lenovo and Acer launching in October, with game publishers including Electronic Arts, Embark and Ubisoft bringing titles to the platform.

Three widely used agent applications are receiving simplified local model setup on Windows. Perplexity Portable Computer, launched last month on Linux for DGX Spark systems with at least 24GB VRAM, is coming to Windows soon. The application lets users run complete workflows locally without consuming cloud credits, with the option to escalate parts of tasks to frontier models when needed. Use cases include engineering teams reviewing GitHub pull requests and fixing out-of-sync documentation, finance professionals analyzing brokerage statements and tax returns locally to identify fee and tax optimization opportunities, and startups analyzing funnel data to identify where user signups drop off between install and first task completion.

Hermes Agent, developed by Nous Research and used by millions, is receiving one-click local setup on Windows. The agent automatically detects the NVIDIA GPU, selects an appropriate model and configuration, and runs it through integrated llama.cpp with NVIDIA optimizations already configured. OpenClaw, the largest AI project on GitHub with over 380,000 stars, is also getting simplified Windows setup for RTX GPUs with at least 24GB VRAM.

Performance improvements include llama.cpp delivering up to 1.9x higher throughput on GeForce RTX 5090 through kernel optimizations and enhanced speculative decoding, with vLLM delivering 1.2x improvements on RTX PRO 6000 and up to 1.4x on DGX Spark clusters. NVIDIA PAIR, a free open-source Personal AI Router, distributes inference requests across PCs on a local network to optimize performance when running multiple independent tasks in parallel.

Read the original