Jiushi CEO Kong Qi and CTO Zhuang Li announced the completion of the industry's first L4-level 10,000-GPU cluster totaling nearly 15,000 cards to support its city-level physical AI upgrade.
On September 10, Jiushi CEO Kong Qi announced a strategic upgrade to “city-level physical AI” in Guangzhou. A week later, on September 17, CTO Zhuang Li revealed the underlying infrastructure powering this shift: Jiushi has built the industry’s first L4-grade 10,000-accelerator-card cluster, with a total compute capacity nearing 15,000 cards. This makes Jiushi the first L4 autonomous driving team to achieve a 10,000-card cluster. Behind this massive compute power lies Jiushi’s APEX multimodal foundation model, which is scaling from the tens of billions to the hundreds of billions of parameters.
The autonomous driving sector has long harbored a persistent bias: passenger autonomy is high-tech, while urban freight robots face a lower AI barrier. Reality proves otherwise. Founded in August 2021, Jiushi’s founding team consists of pioneers in China’s autonomous driving space who previously led Robotaxi and Robotruck R&D at Baidu’s Silicon Valley Research Institute and made significant contributions to the first-generation Apollo open-source platform. From its inception, the company focused on the novel RoboVan domain. Over the past five years, Jiushi has deeply cultivated three core areas: technology, scenarios, and business models.
In these five years, Jiushi accomplished three key milestones: first, delivering tangible technology by inventing the world’s first RoboVan; second, expanding autonomous driving into diverse real-world applications, leveraging L4 technology across industries beyond logistics; and third, establishing a mature commercial framework centered on vehicle sales, services, and capacity provision. Since selling its first vehicle in 2023, Jiushi’s global fleet has exceeded 30,000 units, operating in over 20 countries and 300 cities, with cumulative real-world operational mileage surpassing 270 million kilometers.
However, once scale was proven, a more fundamental question emerged: what separates leaders when autonomous vehicles encounter the vast, rule-defying long-tail scenarios of real cities—unmarked rural roads, temporary construction barriers, irregular obstacles, chaotic human-vehicle mixing, and sudden lighting changes during severe weather? Jiushi now answers this from a new dimension by disclosing its compute cluster has surpassed the “10,000-card” threshold, cementing its position as the industry’s first L4 enterprise to build such infrastructure. This marks the autonomous driving sector’s transition into a new paradigm driven by multimodal foundation models.
A 10,000-card cluster refers to a high-performance computing system comprising over 10,000 accelerator cards designed to train foundational models. While a single AI accelerator card functions as a high-speed computational unit, a 10,000-card cluster operates like an “AI super factory,” breaking down complex model training tasks into parallel processes. Crucially, compute tiering fundamentally dictates the intelligence ceiling of any autonomous driving enterprise. Real-world urban freight scenarios are saturated with long-tail, non-standard road conditions that cannot be exhaustively coded via manual rules. Jiushi’s solution is to use foundation models to learn the physical world, with the 10,000-card cluster serving as the hard prerequisite for training them. It is the robust foundation supporting the evolution of Jiushi’s hundred-billion-parameter model.
According to Jiushi, its APEX multimodal foundation model is advancing from tens of billions toward hundreds of billions of parameters. APEX integrates massive real-world L4 driving data with internet-based language and video knowledge corpora, alongside real operational human decision data. Once base training is complete, APEX empowers downstream data, model, and evaluation pipelines, creating a closed-loop iteration system. The 10,000-card cluster grants Jiushi “self-evolution” capabilities. Daily, 30,000 vehicles across 300+ cities and 270 million kilometers of real L4 operational data feed this compute infrastructure. Stronger compute accelerates model iteration; stronger models reduce operational costs, attract more clients, and generate even more scarce data through increased deployment. As Kong Qi stated, “Complex urban open roads generate high-value data, and the business model itself must be profitable to continuously acquire data and drive technological progress.” Compute is not a cost; it is the starting point of the flywheel.
The core of Jiushi’s strategic upgrade rests on the “Jiushi Brain”—the foundational pillar enabling its autonomous vehicles to serve both commercial operations and urban governance ecosystems. At its heart is the APEX model scaling into the hundreds of billions of parameters, which equips Jiushi with the ability to “understand changes in the physical world.” This capability drives four qualitative shifts. First, enhanced scenario reasoning and decision-making with predictive simulation, significantly improving ride smoothness and safety through long-horizon strategic anticipation. Second, superior knowledge fusion that integrates web videos, traffic texts, and human driving experience to uniformly interpret traffic lights, police gestures, road signs, and local traffic rules, enabling stronger cross-city generalization—the very confidence behind Jiushi’s mapless L4 deployment and full-scenario adaptability on unfamiliar roads. Third, advanced simulation capabilities that highly replicate offline dynamics and human-vehicle behaviors, ensuring credible parallel simulation across millions of scenarios. Numerous edge-case hazards can be verified in the cloud, reducing real-vehicle testing risks. Fourth, higher value in model distillation and edge deployment. A one-stop distillation process creates a complete end-to-end autonomous driving brain with unified weights for perception, prediction, and planning. A single base model adapts to multiple RoboVan variants, enabling batch lightweighting and drastically cutting multi-model iteration costs.
Built upon the APEX multimodal foundation model, Jiushi has engineered a vehicle-cloud synergy mechanism. The logic is straightforward: the more complex the scenario, the deeper the cloud model’s intervention. While current mainstream edge end-to-end models effectively cover general road driving above 30 km/h, they hit capability boundaries. Jiushi’s vehicle-cloud linked architecture deploys models in layers based on road conditions and speed ranges. For complex roads at 5–30 km/h, the system relies on edge Vision-Language-Action (VLA) models for local real-time inference while simultaneously calling cloud VLA models for understanding enhancement. In 0–5 km/h terminal scenarios, beyond the edge and cloud VLA models, Jiushi innovatively introduced a Safety Officer Agent capable of autonomously executing rescue maneuvers in extreme conditions, making high-level decisions such as detouring, rerouting, or pulling over. By combining edge base compute with cloud-enhanced capabilities, Jiushi transcends traditional L2 limitations, achieving full-scenario coverage across normal roads, complex conditions, and emergency rescue.
The essential difference between Jiushi and L2 systems is not speed, but agency: the vehicle recognizes when it “cannot cope,” proactively decelerates or stops, requests cloud analysis, and executes instructions rather than passively awaiting human takeover. This synergy evolves continuously through closed-loop iteration: real-world data feeds back into the APEX base, base improvements support cloud model upgrades, and cloud models distill optimized versions for the edge. Through knowledge distillation and physical simulation, APEX constructs a model matrix spanning edge VLA, cloud VLA, and the Safety Officer Agent. Scenario data generated by each model continuously flows back to refine the base. Notably, Jiushi’s VLA does not force translation into human language before vehicle control. It functions more like a human’s pre-action “thought”—seeing a vehicle ahead and instinctively deciding to brake and yield. This “thought” is faster, more abstract, and more compressed than verbalizing “I need to brake.” By aligning with human language patterns, Jiushi enables models to reason causality before deciding.
The edge ensures real-time safety, the cloud sets the cognitive ceiling, and the Agent safeguards against extreme scenarios. The 10,000-card cluster acts as the hardware factory, APEX serves as the software brain, and 270 million kilometers of real L4 operational mileage forms the scenario foundation. Together, they forge a four-fold self-reinforcing barrier encompassing data, compute, models, and engineering. Jiushi’s strategic upgrade is underpinned by proven commercial scale: operations in over 20 countries and 300 cities, more than 270 million kilometers of actual L4 mileage, a fleet exceeding 30,000 units, and a 100-fold capacity increase over three years. Yet the true value of these figures lies not in scale itself, but in the 270 million kilometers yielding something general-purpose AI companies cannot access: scarce real-world physical data. Every day, each vehicle encounters police gestures, construction barriers, temporary controls, and mixed traffic on open city streets, providing training fuel for APEX. These scenarios cannot be exhaustively enumerated in labs or fully replicated through simulation. They must come from real urban operations. This is the underlying logic of Jiushi’s evolution from an autonomous vehicle manufacturer to a city-level physical AI provider, extending from “freight transport” to “urban governance” where autonomous vehicles become extensions of municipal management capabilities. General AI companies lack physical-world data; traditional auto companies lack large-scale real operational data. Jiushi possesses both. When 30,000 vehicles daily supply a foundation model with scarce physical-world data, scale ceases to be an endpoint and becomes fuel. Upon achieving large-scale commercialization, Jiushi is, in essence, a foundation model company endowed with uniquely scarce data resources.