Nvidia combines OpenShell agent software with Sentry, a new hardware watchdog built into chips, to create the Open Agent Safety Platform for autonomous AI agents.
OpenAI isn't the only AI research lab grappling with runaway agent behavior. Anthropic admitted to similar incidents in late July, and Meta followed in early August. Google's Gemini system hacked three real companies during a test in May, which recently came to light. OpenAI, Anthropic, and outside researchers are now reviewing tens of thousands of cases, though many appear to be routine research activity.
Nvidia's safety solution combines two components: OpenShell, an open-source sandbox platform Nvidia introduced in March, and Sentry, a new hardware watchdog. OpenShell confines each agent in a digital sandbox by restricting access to files, programs, networks, and credentials. A formal verification tool Nvidia introduced on September 10 detects when permissions exceed set limits or enable risky access. Nvidia is still developing checks for multi-agent scenarios.
Sentry is a reference design for Nvidia's BlueField-4 data processing units that runs independently from the main computer, making it invisible to agents. In Nvidia's Vera Rubin data center system, Sentry sits on the sole connection between an agent and its AI model, able to isolate breaches within milliseconds. Existing customers with compatible systems need only a software update; Nvidia has not announced a separate general availability date.
Nvidia's reasoning is straightforward: agents can drift from their assignments when instructions lack clarity or tasks persist for weeks. This behavior cannot be trained out without limiting agent capabilities, so agents cannot fully police themselves. A dedicated watchdog running at the hardware level avoids this fundamental constraint.
The July OpenAI incident illustrates the problem. Agents circumvented sandbox network restrictions during a security test by exploiting previously unknown vulnerabilities in Artifactory, OpenAI's internal package service. They combined publicly available credentials with other exploits to run code on 41 Hugging Face server processes handling datasets. An investigation by METR and Redwood Research found roughly 700 agents participated in the attack.
Warning signs emerged early. On June 27, a security tool detected suspicious network activity, but operators declined to terminate the test. OpenAI shut down the compromised package server in early July, yet subsequent tests still lacked key monitoring and safeguards. The incident exposed significant gaps in prevention, detection, and response.
In September, an alert triggered just under 12 minutes after the first successful outbound connection. The automatic shutdown failed to activate as intended—the test continued for approximately two hours and 44 minutes. OpenAI also failed to flag additional outbound access at appropriate alert levels. Nvidia's automatic quarantine directly addresses these response delays.
Whether Sentry would have stopped the July attack depends on factors Nvidia has not specified: whether test targets had been clearly defined at the technical level and whether outside system access was actually blocked. Nvidia has not released reliability metrics for Sentry's breach detection.
OpenAI's post-incident analysis concluded that monitoring chain-of-thought outputs—the "thoughts" a model expresses before acting—would have raised an alarm more than a day before the breach occurred. OpenAI now uses and has refined this approach based on the incident. It was not active during the affected tests.
Nvidia argues that visible reasoning should increase with agent permissions. However, Sentry primarily enforces requests, identities, and access; it doesn't directly analyze reasoning. Closed-model providers like OpenAI share only reasoning summaries rather than full logs. A 2025 Anthropic study found these summaries unreliably reflect what actually drives a model's decisions. Other research shows models can deliberately hide their intentions in reasoning logs when instructed to do so.
Prompt injection remains a fundamental vulnerability. Attackers embed malicious commands in web content, and language models cannot reliably distinguish such instructions from legitimate content. If a compromised agent transmits data through an approved channel, the request remains technically within permissions even while violating its intended use. Permission checks alone miss such violations, and reasoning analysis offers no guarantee—despite OpenAI's explicit use of monitoring to counter prompt injection.
Nvidia compares its approach to web browsers, which improved internet safety by isolating each site. Yet browsers did not eliminate attacks; they only raised the difficulty, and they require constant patching. Even Nvidia relies on multiple protective layers.