Quanjing Technology는 Moore Threads와 제휴하여 국제적 최첨단 실리콘 대비 낮은 비용으로 생산 등급의 AI 토큰 성능을 제공하는 국내 이종 컴퓨팅 솔루션을 도입했다.
On September 3, Qujing Technology and Moore Threads signed a strategic cooperation agreement. The partnership is built on the deep integration of Qujing Technology’s self-developed domestic PD heterogeneous technology with Moore Threads’ MTT S5000 AI computing cards and the MUSA software platform. Both companies will continue to advance the construction of a domestic, high-quality AI token factory based on Qujing Technology’s ATaaS platform, while jointly developing and promoting integrated solutions centered around the Token Pod (Token Super Node).
The joint solution has already entered production and is handling real token traffic from official workloads of leading model developers. Under equivalent production service standards, the solution balances production-grade performance with unit AI token cost advantages, delivering overall cost-effectiveness that surpasses international advanced computing solutions. This milestone marks the first deployment of domestic PD heterogeneous inference into a live production environment, spearheaded by Qujing Technology and Moore Threads.
Liu Xianhe, General Manager of Qujing Technology’s Token Business Division, and Fei Jingran, Deputy General Manager of Moore Threads’ Strategic Sales Division, signed the agreement on behalf of their respective companies. Also present at the signing ceremony were Wu Yongwei, Professor at Tsinghua University and Chief Scientist at Qujing Technology; Ai Zhiyuan, Founder and CEO of Qujing Technology; Wu Wenjie, President and CFO; Xie Weiyu, Chief Architect; Zhang Jianzhong, Founder, Chairman, and CEO of Moore Threads; along with Moore Threads Senior Vice Presidents Dong Longfei and Luo Wenyong.
As large language model applications enter the stage of scaled commercial deployment, infrastructure competition is shifting from single-card performance to system-level token production efficiency. The key metric for a solution’s commercial value is no longer just whether hardware can run a model, but how much it costs to produce each high-quality AI token under defined service targets.
The cost advantage of the joint solution does not stem merely from hardware pricing differences, but from targeted resource optimization across the Prefill and Decode stages. Moore Threads’ MTT S5000 handles input computation and KV cache generation during the Prefill phase, working in tandem with high-bandwidth GPUs optimized for token generation in the Decode phase.
Leveraging its dense AI compute capabilities, native FP8 precision, and the comprehensive MUSA software stack with robust ecosystem compatibility, the MTT S5000 enables Prefill workloads to be decoupled from traditional inference pipelines. This allows Prefill resources to operate as independently provisioned, billed, and optimized compute pools. Qujing Technology applies its proprietary domestic PD heterogeneous technology to deeply optimize models and coordinate cross-device workflows, ultimately consolidating diverse compute resources into a unified, stable, and highly cost-effective end-to-end inference service.
This architecture reduces Prefill-stage consumption of high-cost resources and eliminates the need to provision entire infrastructure clusters based on peak demands for a single stage. Under current project standards, a Prefill resource pool comprising four to five MTT S5000 units delivers input token processing cost-effectiveness that exceeds international advanced computing solutions.
This demonstrates that the competitiveness of domestic computing power is expanding beyond single-card benchmarks to encompass system-level token production efficiency. Through deep synergy between domestic hardware and PD heterogeneous technology, Qujing Technology and Moore Threads have established a domestic computing pathway for scaled inference of trillion-parameter models that successfully balances production-grade performance with cost efficiency.
Ultimately, cost-effectiveness must be validated by real-world operations. Production environments demand resilience against high concurrency, ultra-long context windows, and continuous traffic fluctuations, placing higher requirements on domestic GPU memory management, cross-device coordination, and service stability.
Production data confirms that the coordinated operation of Moore Threads’ MTT S5000 and Qujing Technology’s domestic PD heterogeneous solution meets complex business scenarios for leading domestic large models. The system achieves an average generation speed exceeding 50 TPS, a KV cache hit rate above 90%, 99.9% operational stability, and production-grade low Time To First Token (TTFT) latency, effectively transforming domestic computing power into scalable, high-quality AI token production capability.
Following multiple rounds of performance iteration and specialized optimization, the MTT S5000 has successfully moved past model compatibility and testing phases to officially join the high-quality AI token production pipeline.
These production achievements exemplify Qujing Technology’s “few models, deep optimization” technical strategy. By continuously refining core models, Qujing Technology converts the hardware capabilities of domestic GPUs into stable, deliverable AI token production capacity, creating reusable manufacturing frameworks.
The “few models, deep optimization” approach focuses engineering efforts on key models with clear scaling demands, driving continuous refinement across models, chips, inference systems, and real-world workload patterns. Success is measured by the ability to simultaneously maintain latency, throughput, stability, precision, long-context handling, and cost efficiency while consistently delivering high-quality AI tokens.
This philosophy guided the design and deployment of the domestic PD heterogeneous solution. Qujing Technology structured the PD-split heterogeneous inference architecture according to the distinct task profiles of Prefill and Decode stages, collaborating with Moore Threads to deeply optimize the MTT S5000 alongside target models, inference systems, and runtime environments. Continuous validation against live traffic allowed both teams to iterate model configurations and service strategies, translating experimental parameter gains into stable, high-quality token output and further proving the viability of “few models, deep optimization” in domestic heterogeneous production scenarios.
The resulting model configurations, traffic characteristics, and operational strategies are systematically consolidated into the ATaaS platform, establishing a reusable production capability framework.
The central goal of this strategic partnership is to leverage both companies’ strengths in inference systems, production operations, and domestic GPU hardware/software platforms to jointly build a domestic AI token factory capable of handling real business workloads and continuously producing high-quality AI tokens.
To achieve this, both parties will launch a hardware-software integrated token production solution based on the Token Pod (Token Super Node). Moore Threads will supply the MTT S5000 AI computing cards and MUSA software platform, while Qujing Technology will use the Token Pod as the delivery vehicle. It will deeply integrate these components with Qujing’s proprietary domestic PD heterogeneous inference system, ATaaS platform, and operational services to form a dedicated inference cluster covering infrastructure, inference optimization, platform management, and continuous operations.
To transition the joint solution from pilot deployments to standardized delivery, both companies will embed lessons learned from domestic PD heterogeneous projects—including model adaptation, cross-device coordination, and production operations—directly into the Token Pod. Featuring shared KV cache, traffic-aware configuration, elastic scaling, and fault isolation, the Token Pod can accommodate mixed domestic and international hardware combinations. It forms a configurable, expandable token production unit tailored to specific business needs, accelerating the deployment of the joint solution across concrete projects.
The continued replication of these solutions rests on Qujing Technology’s existing production foundation. By August 2026, Qujing Technology had completed co-development projects achieving daily production of trillions of high-quality AI tokens, while multiple projects have established daily capacities in the hundreds of billions. This provides valuable operational experience for building and managing larger-scale domestic high-quality AI token factories.
As a pioneer in domestic heterogeneous mixed inference, Qujing Technology has advanced the deployment of domestic PD heterogeneous solutions using production-grade standards. While ensuring functional reliability and precision accuracy, the company continues to drive deep engine-level performance optimizations, ensuring domestic GPU hardware capabilities are genuinely converted into high-quality AI token output.
Looking ahead, Qujing Technology and Moore Threads will further deepen technical collaboration, standard development, and ecosystem partnerships. They plan to expand cooperation into key national sectors including internet enterprises, cutting-edge foundational model companies, and telecommunications operators, driving domestic PD heterogeneous inference into broader production application scenarios.