Microsoft의 맞춤형 Maia 300 가속기가 AI 추론 워크로드에서 Nvidia의 현재 지배적 지위와 비교 평가되고 있다.
Microsoft is preparing to unveil its next-generation Maia 300 AI accelerator as early as September, marking another milestone in its effort to build proprietary AI infrastructure and reduce reliance on NVIDIA processors. The announcement aligns with surging demand for computing power driven by the rapid expansion of generative AI, cloud services, and AI agents.
According to sources familiar with the matter, Microsoft is negotiating with Taiwan Semiconductor Manufacturing Co. (TSMC) to secure production capacity for more than 300,000 Maia 300 chips by 2027. The company could eventually scale the program to exceed one million units. While Microsoft has not officially confirmed these production targets, it stated that its custom silicon initiative is being executed at significantly large volumes.
Microsoft launched its first Maia accelerator in 2023 and introduced the Maia 200 in January 2026. Current Maia 200 production has reportedly remained in the tens of thousands, positioning the Maia 300 as a critical step toward substantially larger-scale deployment.
The new accelerator is designed to handle large AI workloads across Microsoft Azure, with a pronounced emphasis on AI inference. Microsoft intends for its custom silicon to manage traffic from both its proprietary AI services and OpenAI-developed models, aiming to reduce inference costs at Azure scale. However, the company has withheld final specifications, performance metrics, benchmark data, memory capacity, and details regarding the manufacturing process.
For context, the Maia 200 is fabricated using a 3-nanometre process and features 216GB of HBM3e memory, 7TB/s of memory bandwidth, and 272MB of on-die SRAM. Its networking architecture supports clusters of up to 6,144 accelerators. Microsoft reports that the Maia 200 delivers over 10 petaFLOPS at FP4 precision and more than 5 petaFLOPS at FP8 precision, alongside a claimed 30% improvement in performance per dollar compared to the latest hardware in its current fleet.
Microsoft does not intend to replace NVIDIA across all AI workloads. Instead, the Maia series will target tasks that can be optimized for Microsoft’s proprietary software and cloud infrastructure, while NVIDIA accelerators will continue to support workloads where they hold a competitive edge. Additionally, Microsoft claims the Maia 200 offers 40% better performance per watt for Maia models and natively supports both OpenAI and Microsoft AI workloads. The company is also actively courting major cloud customers, including Anthropic, to adopt Maia-based solutions.
The commercial success of the Maia 300 will ultimately hinge on its real-world performance, cost structure, energy efficiency, software compatibility, and supply availability. Until Microsoft discloses further details on its architecture, manufacturing node, memory configuration, benchmark results, and pricing, many questions regarding the chip’s capabilities remain open.