Qualcomm says AI companies want phones running 100-billion-parameter models by 2028.
Qualcomm has received a specific request from several artificial intelligence (AI) companies: phones able to run a model of at least 100 billion parameters continuously by 2028. The account comes from CEO Cristiano Amon, who named none of the customers and promised neither a chip nor a product date.
The remark came in an episode of Sources, the podcast by journalist Alex Heath, published on October 8, 2026 and recorded live at Snapdragon Summit in Hawaii, the event where Qualcomm shows its mobile chips. In the episode summary, Heath writes that AI companies are already asking Qualcomm to support models of at least 100 billion parameters on phones by 2028, with the AI running all day. From there the story spread widely, including through two Reddit threads that carried it to readers of local models.
The source text is short, and it is worth pinning down before commenting on it. Amon is relaying a request from Qualcomm's customers, not announcing a product. Those customers are "some companies at the frontier of AI," unnamed. The deadline is 2028, the requirement is a model of at least 100 billion parameters, and the mode is "always on."
The same episode places the remark inside a wider argument. Amon holds that AI agents, meaning programs that use apps on our behalf and act on the user's personal context, will give the phone a new reason to be bought, and that this shift will tap demand that has stalled because users keep postponing upgrades, partly because of memory prices that have skyrocketed. Heath also points out that Qualcomm supplies the chips in Android phones, in Meta's smart glasses and in future devices from AI companies such as OpenAI, so it has a direct stake in several scenarios for what the next gadget in your pocket will be.
Some outlets have stretched the story. Startup Fortune describes a model that stays "running continuously" rather than being loaded for a task and shut down, and adds that many of these companies are reportedly already building phones. Those are readings of the story, not Amon's words, and the Sources summary does not confirm them. The Reddit headline saying "continuously" captures the gist fairly, but it remains a paraphrase.
**What sits inside a flagship phone today**
On September 22 Qualcomm introduced two chips, the Snapdragon 8 Elite Gen 6 and the 8 Elite Extreme Gen 6, aimed at high-end Android phones. The official brief for the Extreme version lists up to 24 GB of memory and an NPU shared memory (the part of the chip dedicated to AI calculations) that is 50% larger. The standard Gen 6 brief reports the same 24 GB ceiling.
Press coverage adds a figure that appears in neither brief. According to TokenPost, the Extreme Gen 6 is built to run mixture-of-experts (MoE, an architecture where only a share of the parameters works for each generated word) models beyond 30 billion parameters locally. If the figure is right, Amon's target is more than triple what today's silicon promises. TokenPost specifies that Qualcomm has announced no chip able to sustain 100 billion parameters continuously.
Comparing with what is already on the market helps with scale. Google brought Gemini Nano 4 to seven phones, raising the hardware bar for built-in models, and the most discussed open models for PCs, such as Qwen3.8-27B, which runs locally on 16 GB, stay under 30 billion parameters. A 100-billion-parameter model is a different proposition.
**The memory math**
The number that decides everything is memory. A parameter, once compressed with 4-bit quantization (the technique that shrinks each model value from 16 to 4 bits, at the cost of some precision), takes half a byte. One hundred billion parameters at half a byte make 50 billion bytes, so about 50 GB for the weights alone, the model's actual content. At 8 bits it reaches 100 GB. This calculation is ours and leaves out the context token cache (KV cache, the temporary memory that grows with the length of the conversation) as well as the operating system.
On the other side sits the 24 GB ceiling from the briefs. At 4 bits, 24 GB holds at most 48 billion parameters, and in practice far fewer, because Android, open apps and the cache take their share. The gap against the 50 GB needed is a factor of two even in the best case, and a phone cannot hand all its RAM to a single program.
That is why attention turns to MoE models and to more efficient use of memory. In a MoE only a fraction of the parameters is active for each token, so the compute required drops, but all the parameters still have to live somewhere, in RAM or in the phone's flash storage (roomier and slower). Moving unused weights onto flash is one route, and that is where the next two years will be decided, between read speed, power draw and overheating.
Users on r/LocalLLaMA, the community that discusses models to run locally, did the same math with equal skepticism. The thread gathered almost 600 points and hundreds of comments, with some noting that a 27-billion-parameter model already strains a laptop, others joking about the RAM left for a context window of a few thousand tokens, and others imagining 64 GB phones. A parallel thread on r/singularity gathered about 180 points.
**Why memory is the bottleneck, not the chip**
The problem is not only technical. In the same conversation Amon cites memory prices as a brake on purchases, and the supply picture does not help. A few days ago we reported that for Micron, memory will be scarcer in 2027 and 2028 than in 2026, which is the very window in which Qualcomm's customers would like their 100-billion-parameter phones. Anyone designing a device with 48 or 64 GB of RAM has to buy it at prices nobody dares to estimate two years out.
The same tension shows up in the PC world. Microsoft has just announced DeepSeek V4 Flash and Nemotron running locally on Windows, and on the graphics card front the RTX 5090 is reportedly out of production according to two leakers, with a 24 GB RTX 5080 taking the role of new flagship GeForce. In both cases the number being cited is the memory available for the weights, not the teraflops.
**The market's point of view**
There is a commercial reading that should not be ignored. A chief executive who talks about 100-billion-parameter phones on a podcast, in a year when the sector lives on agents and promises about the next upgrade cycle, is describing a purchase cycle, and a purchase cycle is worth more than a chip. The 100-billion figure is convenient: big enough to make the device in your pocket look obsolete, and far enough out (2028) that no test can contradict it today.
None of this makes the request false. AI companies have concrete reasons to want models on the phone. Personal data nobody wants to send to a data center stays local, the answer arrives even without signal, and above all inference costs (the compute behind each answer) move from the provider's bill to the user's battery. For a company serving hundreds of millions of people, that is a far stronger argument than data sovereignty.
**What is left to verify**
Three things remain open. The first is the names. Amon did not say who asked for what, and the fact that Qualcomm supplies chips to OpenAI does not mean the request comes from there. The second is the model format. A 100-billion-parameter MoE with a few billion active parameters and a dense 100-billion-parameter model are two different problems for memory and power. The third is what "continuous" means. An assistant that keeps the model loaded in the background and wakes it on request is feasible with flash. One that generates permanently, as "all day" suggests, would drain the battery in a few hours.
In languages other than English and Italian I looked for independent confirmation of the figure (French, German, Japanese, Chinese) and found none. That is no sign of a hoax, since the primary source exists and is the Sources episode, but it means everything we know today passes through a single interview.