Saturday, October 10, 2026
AI Infrastructure · News & Analysis
Home › Chips & Hardware › Report
Chips & Hardware · Report

Nvidia has released a fix for a high-severity flaw, tracked as CVE-2026-47483, that could crash GPU monitoring on exposed servers.

Operators running Nvidia GPU servers should patch promptly, since a crash in monitoring tooling could blind fleet health checks on exposed machines.
Trade pressSlicast · October 8, 2026 at 18:17 UTC · Global · Source: The Register
importance 40

Researchers found thousands of GPU servers exposing Nvidia's DCGM Exporter to the internet. Hundreds of them are potentially vulnerable to a high-severity flaw that could let unauthenticated attackers crash the GPU monitoring service and disrupt AI workloads.

DCGM Exporters read telemetry from the GPUs on a host, including hardware details, utilization, memory usage, power consumption, and error events. Each GPU has its own unique ID, or UUID, and all of these metrics are exposed in plaintext over HTTP. This gives would-be attackers detailed information useful for reconnaissance, including mapping GPU infrastructure, identifying potentially vulnerable systems, and monitoring workload activity.

Michael Katchinskiy, a researcher at datacenter security startup Lava, found and reported the bug in the GPU health and performance monitoring service. In September, GPU maker Nvidia released a fix for the flaw, tracked as CVE-2026-47483, and assigned it an 8.2 CVSS high-severity rating.

"Once we realized how much these endpoints revealed, the next question was: How many of them are exposed to the internet?" Katchinskiy said in a blog post on Thursday.

The researchers began scanning the internet for exposed DCGM Exporters, and the scale of the exposure proved "especially significant," he wrote. Over four scans between March and May, the threat hunters found about 2,100 GPU servers exposing DCGM Exporter metrics to the open internet. Those servers revealed 12,000 GPU UUIDs, and none of the endpoints required authentication. The hosts belonged to about 300 organizations, according to Katchinskiy. Nearly half of the exposed GPUs, 5,274 or 44 percent of the total, were located in the US.

These GPUs represented about $100 million in hardware. They included Nvidia Blackwell Ultra B300, H200, and H100 GPUs, which are used to run large-scale AI workloads, as well as consumer RTX 5090 and 4090 systems.

While investigating the exposed systems, the Lava team found that about 25 percent of the exposed DCGM hosts also revealed data from Go's /debug/pprof/ built-in profiling tool. The profiler exposes runtime performance data such as CPU usage, memory allocations, goroutine states, and blocking events for running Go applications.

"With enough concurrent unauthenticated requests, the exporter could run out of memory and crash, cutting off visibility into GPU health and activity," Katchinskiy wrote. The resulting CPU and memory pressure could also affect AI training or inference workloads.

Nvidia fixed the issue in version 4.8.2, and operators should upgrade to that version or later.

Lava also examined Prometheus Node Exporter, which monitors server hardware and operating systems and also exposes metrics over HTTP. The team found 12,096 public Node Exporter hosts exposing data on server models, operating systems, firmware versions, hostnames, storage paths, and networking hardware commonly used in GPU clusters. This information reveals how environments are built and configured, and attackers could use it for reconnaissance by matching systems to known vulnerabilities.

The publicly exposed monitoring services affected customer infrastructure at neocloud and GPU cloud providers including Nebius, Voltage Park, Lambda, Northern Data, and DigitalOcean. Lava reported its findings to the affected providers, and Katchinskiy says the providers worked with customers to address the exposures.

Operators should not allow Nvidia DCGM Exporter, Node Exporter, or Prometheus services to be directly reachable from the public internet. Lava recommends restricting access to authorized monitoring infrastructure.

"The findings highlight a growing security gap in AI infrastructure: companies are spending millions on GPUs while leaving critical systems exposed," Katchinskiy wrote. "Those exposures can reveal how AI environments are built and, in some cases, allow attackers to disrupt them."

Read the original
Nvidia has released a fix for a high-severity… · Slicast