Analysis identifies DeepSeek's competitive success as creating business opportunities for IT infrastructure providers.
Chinese AI company DeepSeek made waves in the technology industry with claims of achieving performance comparable to leading AI models while dramatically reducing training infrastructure requirements. DeepSeek unveiled its DeepSeek-R1 model, which it claims rivals leading AI systems like OpenAI's GPT-4 and Meta's Llama. The significance lies not in the model itself, but in the techniques employed to build it. According to DeepSeek, its model was trained on just 2,048 Nvidia H800 GPUs, costing approximately $5.58 million — a fraction of the infrastructure and cost typically associated with such efforts. By employing advanced techniques such as FP8 precision, modular architecture, and proprietary communication optimizations like DualPipe, DeepSeek has purportedly streamlined AI training to a level previously thought unattainable.
If DeepSeek's claims are validated, they remove the almost insurmountable cost barriers to AI training, opening the door to much broader adoption and competition in the market. One of the challenges of current AI training techniques is that the resources are prohibitive, requiring investment that's only feasible for the largest hyperscalers. DeepSeek's approach promises to disrupt that model by making AI training accessible to most enterprises. Despite reducing reliance on high-end GPUs, DeepSeek's approach does not eliminate the need for robust supporting infrastructure — key requirements like high-performance storage, low-latency networking, and strong data management frameworks remain critical. If DeepSeek's claims hold, rack-level training clusters may now be possible, representing a moment that mirrors historic IT transformations like the transition from mainframes to mini-computers and ultimately PCs.
This shift creates tremendous opportunities for storage providers, server OEMs, and networking companies. Companies like NetApp and Pure Storage and server manufacturers like Dell Technologies, HPE, and Lenovo all stand to benefit as demand for scalable and cost-effective AI infrastructure grows. Nvidia's stock price sharply dropped following DeepSeek's announcement, but the company remains well-positioned due to its entrenched ecosystem, which includes its CUDA platform and investments in AI systems like DGX and Mellanox networking solutions. In its most recent earnings, Nvidia reported that networking revenue increased 20% year over year, with significant growth in Spectrum-X, up 3x year over year. Software revenue is annualizing at $1.5 billion, about 4% of its total revenue, and the company expects that to exceed $2 billion by year-end, driven by offerings like NVIDIA AI Enterprise, Omniverse, and AI microservices.
The stock market also reacted negatively to custom silicon companies like Broadcom and Marvell, yet DeepSeek's approach may prove to be a net positive for these companies. DeepSeek's announcement, focused on AI training, should have minimal impact on inference work or custom Arm-based processors for cloud service providers. The demand for low-latency, high-throughput networking solutions remains essential in DeepSeek's framework. Broadcom's dominance in Ethernet and InfiniBand and Marvell's strength in energy-efficient and high-bandwidth interconnects position both companies to benefit from the need for advanced interconnects in decentralized AI training environments. For both companies, DeepSeek's innovation represents less of a disruption and more of a realignment of market dynamics. If reproducible, DeepSeek's claims will drive a shift in the AI training landscape by lowering costs and democratizing access to advanced model training.