Nvidia announces Turing-based GPUs for AI inference applications.
Machine learning's evolution has shifted focus from neural network training toward inference—the act of using trained models on new data to create useful applications. Nvidia, while dominant in training infrastructure, has been developing inference solutions since at least November 2015, when it released the Tesla M4 and M40 accelerators based on Maxwell GPUs. The Tesla M4 delivered 2.2 teraflops of single-precision performance with 1,024 CUDA cores and 8 GB of GDDR5 memory providing 88 GB/sec of memory bandwidth, available in 50-watt and 75-watt thermal envelopes. Though this represented significant improvement over CPU inference, datacenters predominantly kept inference workloads on existing Xeon servers.
Two years ago, Nvidia's launch of the Tesla P4 and P40 accelerators marked a more serious push into inference. This timing also saw the introduction of what became known as Buck's Law, a concept articulated by Ian Buck, vice president and general manager of Nvidia's Tesla datacenter business unit. Buck's Law posited that "every bit of data created in the world would require a gigaflops of compute over its lifetime"—an observation reflecting the scale of computation required across hyperscalers and enterprises. Facebook processes over 1 billion videos viewed per day through machine learning recommendation engines, Google and Bing handle over 1 billion searches driven by speech recognition daily, and these activities collectively drive over 1 trillion advertising impressions per day, all demanding efficient inference infrastructure.
The Tesla P4 significantly advanced GPU inference capabilities with 2,560 cores, delivering 5.5 teraflops of single-precision performance and 22 teraops using the industry-standard INT8 eight-bit integer format. The accelerator featured 8 GB of GDDR5 memory with memory bandwidth of 192 GB/sec—more than double the Tesla M4's bandwidth to match its more than doubled compute capacity. According to Buck, while he could not disclose specific P4 revenues, inference workloads broke down as follows: video recognition represented half the business, speech processing about a quarter, and search constituted another significant portion, with inference revenue having become "a material part of the business."
Nvidia's announcement of the Tesla T4 accelerator at the GPU Technical Conference in Japan represented the next generational shift in inference acceleration. This advancement arrives as machine learning inference opportunities accelerate, with the addressable market representing $20 billion in infrastructure sales over the next five years. Notably, just two years prior to the T4's introduction, 95 percent of machine learning inference still ran on standard X86 servers rather than specialized accelerators, but the volume of data requiring efficient inference processing—whether on edge devices, in-network systems, or datacenter applications—now demands fundamentally more efficient approaches.