Qualcomm secured its first major Western hyperscaler contract with AWS, valued at up to $60 billion for inference silicon.
Amazon Web Services has become Qualcomm’s first Western hyperscaler customer for custom AI inference silicon, signing a multi-generational co-development agreement that includes a warrant allowing Amazon to acquire up to 25 million Qualcomm shares. The full block vests only if Amazon commits up to $60 billion toward Qualcomm chips, networking hardware, and manufacturing services through September 2036. The arrangement was disclosed in Qualcomm’s SEC Form 8-K, filed on September 8, 2026. Qualcomm CFO Akash Palkhiwala clarified at the Goldman Sachs Communacopia & Technology Conference on the same day that the partnership is not a forward-looking promise; revenue will begin in Qualcomm’s fiscal first quarter of 2027—the December 2026 quarter—with chips already in production.
Following the announcement, QCOM shares surged up to 9.5% during Tuesday’s trading session, reaching an intraday high of $183.49. This marked the stock’s strongest single-day gain in years and its first positive return for the year. Conversely, Amazon shares dipped approximately 1% as investors assessed the integration of a new silicon supplier into the AWS infrastructure stack.
The partnership spans two distinct product tracks. The first focuses on customized AI inference silicon co-designed across multiple chip generations. The second involves a joint development effort for optical connectivity hardware capable of scaling to 1.6 terabits per second (1.6T), designed to support intra-cluster networking—the critical infrastructure that routes data between server racks within hyperscale data centers.
**What the Deal Actually Covers — and Why LPDDR Matters**
Both tracks reflect a deliberate architectural divergence from competitors: Qualcomm utilizes LPDDR memory instead of High Bandwidth Memory (HBM), which currently powers Nvidia’s B200 accelerators and AMD’s Instinct MI350X cards. Per Qualcomm’s official specifications, Nvidia’s flagship B200 contains approximately 180 gigabytes (GB) of HBM3e per card, whereas Qualcomm’s AI200 supports up to 768 GB of LPDDR—more than four times the capacity. While HBM provides significantly higher peak memory bandwidth and dominates training workloads requiring simultaneous parallel matrix operations, inference operates differently. Specifically, the “decode” phase, where an AI model generates tokens sequentially, is not primarily bandwidth-bound. Instead, it is memory-capacity-bound; the critical constraint is whether the complete model can remain resident on the accelerator to avoid costly data transfers during user requests.
Qualcomm’s strategy hinges on the premise that decode workloads prioritize capacity over peak bandwidth, and that LPDDR provides adequate throughput at a substantially lower cost per gigabyte than HBM. This approach is further advanced by the AI250, Qualcomm’s next-generation accelerator slated for calendar 2027. The AI250 employs a “near-memory computing” architecture that physically stacks compute logic adjacent to memory, drastically reducing data travel distance. According to Qualcomm’s specifications, this design yields more than 10 times the effective memory bandwidth of conventional DRAM configurations while consuming significantly less power.
At the Goldman Sachs conference, Palkhiwala elaborated on the architecture: “We have this innovative technology that we call High Bandwidth Compute, which is really kind of stacking compute and memory together to deliver extremely high bandwidth that our customers see as a perfect solution for certain kind of decode workloads within inference.” He noted that the technology will eventually expand to cover prefill workloads and the broader inference stack.
**The Optical Layer: Why Qualcomm Bought Alphawave**
The 1.6T optical connectivity component stems directly from Qualcomm’s $2.4 billion acquisition of Alphawave Semi, announced in June 2025 and finalized around December 2025. Alphawave specializes in high-speed serializer/deserializer (SerDes) intellectual property—the underlying technology that converts parallel internal chip data into the serial format required for ultra-fast external communication.
Alphawave’s portfolio encompasses PCIe Gen 6, CXL 3.0, and Ethernet IP rated for 400G, 800G, and 1.6T data rates, alongside chiplet interconnect standards like UCIe. In March 2026, Qualcomm and Lightmatter jointly demonstrated 1.6 Tbps per-fiber optical throughput utilizing a 16-wavelength dense-wavelength-division multiplexing (DWDM) architecture. The company claims this design achieves eight times greater bandwidth density per fiber compared to competing near-packaged optics solutions.
Integrating inference compute with optical transport within a single partnership addresses a critical scale-out challenge: as AI models expand, network fabric bandwidth connecting data center racks becomes a primary bottleneck independent of compute silicon. A vendor controlling both the inference accelerator and the data transport layer can optimize the entire data path end-to-end, eliminating the friction of integrating disparate third-party products. Palkhiwala highlighted this synergy at the Goldman Sachs conference: “Even our custom silicon engagement includes Alphawave SerDes, and they have optical connectivity products that are also a part of the Amazon agreement.”
**What This Deal Is — and What It Isn't**
AWS currently operates two proprietary AI accelerator programs: Trainium for model training and Inferentia for inference. Industry analysis of Marvell’s co-design role in AWS silicon indicates these programs primarily serve AWS’s internal AI workloads and power the Trn and Inf instance types offered on its public cloud. Qualcomm’s custom silicon occupies a different strategic niche. Rather than displacing Trainium or Inferentia internally, it will be integrated into AWS’s broader cloud infrastructure to host third-party and commercial inference workloads at scale.
This distinction clarifies why AWS partnered with Qualcomm instead of merely expanding its Trainium program. The inference market is increasingly disaggregating by workload subtype. Prefill workloads—which handle the initial processing of user input—and decode workloads—which generate tokens sequentially—demand fundamentally different hardware characteristics. For a hyperscaler operating at Amazon’s scale, routing distinct workload types to specialized silicon architectures presents a clear economic advantage, minimizing per-token costs. Qualcomm’s LPDDR-based design, engineered specifically for decode efficiency and memory capacity, functions as a complementary tier rather than a direct competitor.
Palkhiwala further confirmed at the Goldman Sachs conference that a second, unnamed global hyperscaler is pursuing a partnership with a “similar technology scope” to the AWS agreement, indicating that the deal structure is designed to be replicable rather than a singular transaction.
**The Warrant Structure: Pay-to-Play at Scale**
The equity warrant structure outlined in Qualcomm’s SEC Form 8-K establishes strict performance accountability. Amazon’s affiliate holds the right to purchase up to 25 million QCOM shares at an exercise price of $161.26 per share, with the warrant expiring on September 3, 2036, and cashless exercise provisions permitted. At the strike price, the maximum warrant carries a face value of approximately $4 billion.
Share vesting occurs in tranches directly linked to commercial milestones: the execution of binding agreements, the placement of purchase orders, and the actual procurement of Qualcomm server chips, technology, systems, and manufacturing services. An initial tranche of 3.75 million shares (15% of the total) vested upon Amazon’s initial purchase commitments. The remaining 85% will vest progressively as Amazon scales its spending, capped at the $60 billion threshold over the ten-year period.
This framework operates as a mutual performance guarantee. Amazon’s equity upside materializes only if Qualcomm executes at scale, while Qualcomm’s revenue growth depends entirely on the continued pace of Amazon’s AI infrastructure expansion. Neither party profits from the other’s underperformance. This “pay-to-play” model, which ties warrant vesting to commercial milestones rather than calendar dates, is emerging as a standard template for major hyperscaler chip partnerships—a structural parallel to the warrant Amazon issued against STMicroelectronics in a separate semiconductor supply agreement earlier in 2026.
**The Strategic Context: Decoupling from Apple**
Tuesday’s announcement arrives at a pivotal juncture for Qualcomm. The company derives the bulk of its revenue from smartphone processors and modem IP, with Apple historically representing a substantial share of that business. Apple’s deployment of its proprietary C1 modem—initially launched in the iPhone 16e—marks the beginning of a phased transition that CEO Cristiano Amon has confirmed will fully replace Qualcomm modems by 2027. Consequently, Qualcomm’s share of Apple’s device modem market is projected to decline from approximately 70% in 2025 iPhones to 20% in 2026 models, to z