SoftBank launches Infrinia, an AI Cloud OS for GPU cloud services and AI data centers.
SoftBank has announced that its Infrinia Team, which develops next-generation AI infrastructure architecture and systems, has created "Infrinia AI Cloud OS," a software stack designed for AI data centers. The platform enables AI data center operators to build Kubernetes as a Service (KaaS) in multi-tenant environments and Inference as a Service (Inf-aaS) that provides Large Language Model inference capabilities via APIs as part of their own GPU cloud services. The software stack is expected to reduce total cost of ownership (TCO) and operational burden compared with bespoke solutions or in-house development, enabling rapid delivery of GPU cloud services that efficiently and flexibly support the full AI lifecycle from model training to inference.
The development of "Infrinia AI Cloud OS" addresses rapidly expanding demand for GPU-accelerated AI computing across generative AI, autonomous robotics, simulation, drug discovery, and materials development fields. As user needs and usage patterns for AI computing have become increasingly diverse and sophisticated, GPU cloud service providers have faced complex operational challenges requiring highly specialized expertise. These include managing fully abstracted GPU bare-metal servers, providing cost-optimized inference services without GPU management concerns, and enabling advanced operations where AI models are trained and optimized on centralized servers and deployed for inference at the edge.
"Infrinia AI Cloud OS" addresses these challenges through comprehensive automation and intelligent management features. The platform automates the entire stack—from BIOS and RAID settings to the OS, GPU Drivers, networking, Kubernetes Controllers, and Storage—on state-of-the-art GPU Platforms such as NVIDIA GB200 NVL72. It enables software-defined dynamic, on-the-fly physical connectivity reconfiguration of NVIDIA NVLink and Inter-Node Memory Exchange as customers create, update, and delete clusters. The system automatically allocates nodes based on GPU proximity and NVIDIA NVLink domain to reduce latency and maximize GPU-to-GPU bandwidth, while enabling users to deploy inference services simply by selecting Large Language Models without working directly with Kubernetes or underlying infrastructure.
The platform offers OpenAI-compatible APIs for drop-in integration with existing AI applications and seamless scaling across multiple nodes in core and edge platforms including NVIDIA GB200 NVL72. Security and operability features include tenant isolation through encrypted cluster communications and separation, automation of operational maintenance including system monitoring and failover, and an API environment for connecting to the AI data center's portal, customer management systems, and billing systems. SoftBank plans to deploy "Infrinia AI Cloud OS" initially within its own GPU cloud services, with the Infrinia Team—established within SB Telecom America, a wholly owned subsidiary of SoftBank—aiming to expand deployment to overseas data centers and cloud environments with a view to global adoption. According to Junichi Miyakawa, President & CEO of SoftBank, "To further deepen the utilization of AI as it evolves toward AI agents and Physical AI, SoftBank is launching a new GPU cloud service and software business to provide the essential capabilities required for the large-scale deployment of AI in society... Through Infrinia, SoftBank will play a central role in building the cloud foundation for the AI era and delivering sustainable value to society."