Saturday, August 8, 2026
DarkSubscribe
AI Infrastructure · News & Analysis
HomeChips & HardwareReport
Chips & Hardware · Report

Nvidia announces Turing, a groundbreaking GPU architecture with hardware-accelerated real-time ray tracing for datacenter and design workloads.

Represents major architectural leap enabling new AI/ML workload classes and reinforcing Nvidia technology leadership in accelerated computing.
Trade pressSlicast · August 15, 2018 · Global · Source: top500.org
importance 85

At SIGGRAPH 2018, NVIDIA CEO Jensen Huang unveiled the Turing architecture, a new GPU design that merges AI, ray tracing, rasterization, and computation. Calling it the greatest leap in the graphics processor since the CUDA GPU was introduced in 2006, Huang told the audience the new Turing GPUs will enable designers and artists to render photorealistic scenes in real time. "This fundamentally changes how computer graphics is going to be done," he said. "It's a step function in realism."

The first products incorporating the Turing architecture are the Quadro RTX Professional GPUs, aimed at professionals doing video content creation, and automotive and architectural design, as well as researchers doing scientific visualization. The product line includes the Quadro RTX 8000 with 48GB of memory, 10 Giga-rays/sec ray tracing throughput, 4,608 CUDA Cores and 576 Tensor Cores at an estimated price of $10,000; the Quadro RTX 6000 with 24GB of memory, 10 Giga-rays/sec ray tracing, 4,608 CUDA Cores and 576 Tensor Cores at $6,300; and the Quadro RTX 5000 with 16GB of memory, 6 Giga-rays/sec ray tracing, 3,072 CUDA Cores and 384 Tensor Cores at $2,300. These GPUs will be available in both workstations and servers beginning in the fourth quarter of the year. NVIDIA is also launching the RTX Server based on the top-of-the-line Quadro RTX 8000, designed to serve as the basis of render farms. According to NVIDIA's calculations, four RTX servers equipped with eight GPUs apiece can provide the same rendering throughput as 240 dual-socket servers powered by Skylake CPUs, reducing capital cost from $2 million to $500,000 and power requirements from 144 to 13 kilowatts.

The Turing architecture is built around three specialized hardware elements: a streaming multiprocessor (SM) for compute and shading, Tensor Cores for deep learning and AI, and RT Cores for ray-tracing. The chip contains up to 18.6 billion transistors on 754 mm² of die, nearly the same size as the Tesla V100 GPU, which sports 21.1 billion transistors on 815 mm² of die. The RT Cores are specialized circuitry that enables real-time ray tracing for accurate shadowing, reflections, refractions, and global illumination, allowing the new Quadros to simulate up to 10 billion rays per second. Memory is based on GDDR6, a departure from the previous Quadro GV100 which incorporated 32GB of HBM2 memory, though memory capacity can be effectively doubled by connecting two GPUs via NVLink.

The Turing silicon is backed by NVIDIA RTX, a software stack that includes support for Pixar's Universal Scene Description (USD) and the Material Definition Language (MDL), as well as APIs for rasterization, ray tracing (OptiX, DXR, and Vulkan), simulation (PhysX, FleX and CUDA 10), and AI (NGX SDK). The streaming multiprocessor features separate floating point and integer pipelines that can operate simultaneously, enabling address calculations and numerical calculations at the same time. The new Quadro chips can deliver up to 16 teraflops and 16 teraops of floating point and integer operations, respectively, in parallel, with a unified cache featuring double the bandwidth of the previous generation architecture.

The Turing Tensor Cores support graphics and visualization applications including AI-based denoising, deep learning anti-aliasing (DLAA), frame interpolation, and resolution scaling to reduce render time, increase image resolution, or create special effects. While similar to those in the Volta-based V100 GPU, the new Turing Tensor Cores have significantly boosted tensor calculations for INT8 (8-bit integer) operations, used for inferencing neural networks, from 62.8 teraops in the V100 to 250 teraops in the Quadro RTX chips. The new Tensor Cores also provide INT4 (4-bit integer) capability for certain types of inferencing work requiring less precision, achieving 500 teraops—half a petaop—and supporting 125 teraflops for FP16 data, the same as the V100.

Read the original
Nvidia announces Turing, a groundbreaking GPU… · Slicast