Thursday, September 24, 2026
AI 인프라 · 뉴스 & 분석
반도체·하드웨어리포트
반도체·하드웨어 · 리포트

Alibaba announced a comprehensive full-stack AI strategy roadmap spanning chips, cloud infrastructure, models, and agents.

China's largest cloud provider's vertically integrated chip-to-cloud strategy strengthens domestic AI infrastructure autonomy and directly competes with U.S.-dominated GPU cloud vendors.
업계 전문지Slicast · 2026년 9월 23일 20:42 UTC · 글로벌 · 출처: HPCwire
중요도 75

Alibaba has announced comprehensive updates to its full-stack AI strategy, unveiled at the company's annual Apsara Conference in Hangzhou on September 23, 2026. The roadmap spans innovations in foundation models, multimodal capabilities, proprietary AI chips, agentic cloud infrastructure, and mobile AI agents.

Joe Tsai, Chairman of Alibaba Group, framed the strategic direction: "Over the past few years, AI has continuously expanded our imagination of technological capabilities. AI possesses vast potential for development—it can be deployed and scaled in real-world scenarios, boosting productivity across thousands of industries. This is the true meaning of 'Intelligence Goes Beyond': guiding AI from technological breakthroughs toward value creation."

Eddie Wu, CEO of Alibaba Group, emphasized the scale opportunity ahead. "Today, the total volume of Machine Thinking is less than 3% of all Human Thinking. If that volume eventually scales to 1,000x human capacity, the simple math tells us: Machine Thinking still has an enormous growth runway. As machines are becoming the primary force behind Thinking, turning intelligence into a commodity supplied at scale, the truly groundbreaking products of the Machine Intelligence era have not yet arrived. With this in mind, our target is that by 2032, the global data center capacity operated by Alibaba Cloud will surpass 20GW, fueling the industry's exponentially rising demand for AI."

**Foundation and Multimodal Models**

Alibaba revealed that Qwen 4 is currently in training, with roadmaps announced for Qwen 4.5 and Qwen 5, projected to scale to 5 to 10 trillion parameters. The company has demonstrated progress in Recursive Self-Improvement (RSI) driven by empirical feedback. Over one month of fully automated runs spanning pipeline design, data validation, iterative experimentation, and error diagnosis, Qwen3.8-Max completed 33 iterative cycles, boosting its Artificial Analysis score from 40 to 45 through autonomous training optimization and post-training techniques. In a chip design experiment, the model underwent over 60 hours of self-improvement across the entire design lifecycle, making more than 10,000 EDA tool calls to produce production-grade chip bus modules, reducing chip area by 42% with zero performance compromise.

In multimodal capabilities, Alibaba debuted Qwen3.8-LiveTranslate, a simultaneous interpretation model that reduces latency by nearly 20%—from 2.8 to 2.3 seconds—while enhancing translation fidelity, fluency, and brevity for real-time translation. Qwen-Audio-3.1-TTS-Next, a next-generation audio generation model, creates complete cinematic soundscapes by blending dialogue and ambient sounds from text scripts, designed for professional creative applications including audiobooks, film and television, podcasts, and games. The updated speech suite includes Qwen-Audio-3.1-ASR (Automatic Speech Recognition), Qwen-Audio-3.1-TTS (Text-to-Speech), and Qwen-Audio-3.1-Realtime for enhanced voice AI capabilities.

Qwen-Image 3.1, launching later this year, optimizes image generation for creative design and e-commerce marketing, cutting visual design turnaround from hours to seconds while offering native transparent-background generation and advanced image-editing capabilities. To standardize world model evaluation, Alibaba Token Hub, alongside leading academic and industry partners, introduced a six-tier capability framework and launched the Happyworld Arena benchmarking platform, spanning video generation, embodied AI, and spatial reconstruction.

Alibaba has launched Qwen Intelligence, a business-facing full-stack agentic solution optimized for smartphones, providing phone makers with access to a Qwen-powered agent platform for next-generation AI phones capable of handling complex, cross-app tasks reliably.

**AI Chips and Infrastructure**

T-Head, Alibaba's chip design unit, unveiled the Zhenwu V900, its latest AI training and inference processor delivering three times the performance of its predecessor, the Zhenwu M890 (released in May). Featuring 216 GB of GPU memory and 1,200 GB/s of inter-chip bandwidth, the accelerator is built to handle demanding AI workloads with native support across multiple data precisions, including FP8 and FP4. By reducing inference costs while increasing compute density, it seamlessly handles both high-precision model training and ultra-low-precision inference, scheduled for mass production and commercial release in Q1 2027. T-Head's Zhenwu AI chips have been serving over 650 customers across automobile, finance, large language models, embodied intelligence, energy, and manufacturing.

Alibaba also unveiled an upgraded supernode server integrating the Zhenwu V900 processor, ICN Switch, Panmai SmartNIC, and Zhenyue SSD controller chip, optimized for full-stack system-level synergy across compute, storage, and networking. The server supports supernode clusters comprising up to 500,000 cards.

Additionally, Alibaba unveiled its roadmap for next-generation proprietary CPUs tailored for agentic AI tasks, scheduled for 2027 launch. The Yitian 720 features enhanced single-core performance, higher core density, and increased energy efficiency compared to its predecessor, the Yitian 710. The Yitian 730, the first CPU built on T-Head's proprietary microarchitecture, boasts up to 40% increased SPECint2017/GHz performance over Yitian 710.

**Agentic Cloud Platform**

Alibaba Cloud unveiled a comprehensive suite of upgrades centered on agentic cloud strategy, built around three core scenarios: model, harness, and context.

AI Native Cloud focuses on large-scale model training and inference. Alibaba Cloud's Platform for AI (PAI) integrated synergistic optimizations spanning inferencing, caching, sample replay, and training services. In real-world post-training of Qwen models, PAI completed state-of-the-art model training in just five days. Cloud Parallel File Storage (CPFS), a new-generation system optimized for AI model training, delivers hundred-terabyte-per-second throughput and hundred-million input/output operations per second (IOPS), cutting enterprise AI storage costs by 69%. HPN 8.0 Pro, Alibaba Cloud's latest proprietary AI networking architecture, delivers 100 petabits bandwidth with ultra-low latency. A single cluster supports over 130,000 network ports at 800G speeds, with built-in redundancy ensuring neither optical transceiver nor link failures cause service interruptions, reducing the impact of network upgrades and failures from 50% to 25% compared to previous generations.

Agent Native Cloud enables enterprise-grade agent deployment, operation, and security. AgentCore is a new enterprise platform to build, run, and manage AI agents throughout their entire lifecycle, enabling businesses to build and run their AI tools, safely control human-agent collaboration, and monitor performance for continuous improvements. With enterprise-grade security embedded at the operating layer, it provides organizations a single, controlled foundation to deploy diverse AI agents and seamlessly integrate them into critical business systems at scale. The Agent Security Center delivers full-lifecycle security and compliance management for enterprise AI agent applications through real-time threat detection, ensuring agents operate securely within controllable boundaries.

Context Engine provides real-time data and long-term memory. Agent Context is an enterprise-grade context data service giving AI agents real-time context and long-term memory by connecting documents, business systems, chat records, and multimodal data into a single foundation, enabling agents to remember past tasks, share knowledge across teams, and learn continuously. In knowledge-intensive scenarios such as customer service, AI coding, and data analytics, it cuts token usage by up to 67%. OpenLake, upgraded into a unified multi-modality data lakehouse, serves structured, semi-structured, unstructured, vector, and streaming formats to multiple compute engines for processing, search, analysis, and model training. Compared with traditional architectures, OpenLake reduces total costs by 38% and cuts query response times by 40%.

원문 보기
Alibaba announced a comprehensive full-stack… · Slicast