DeepSeek has raised its API pricing, marking a departure from its previous aggressive discounting strategy amid tightening global compute resource availability.
DeepSeek has introduced a peak-and-off-peak pricing mechanism for its API services, marking a strategic shift from its long-standing reputation as the “price butcher” of the large language model industry. Under the new structure, peak hours are defined as 9:00 to 12:00 and 14:00 to 18:00 Beijing Time, with all other periods classified as off-peak. For example, DeepSeek-V4-Pro’s off-peak rates stand at 0.15 yuan per million tokens for cached inputs, 4.5 yuan for uncached inputs, and 13.5 yuan for outputs. These fees will double during peak hours. The company stated that the adjustment aims to allocate resources more efficiently and incentivize users to schedule tasks according to actual demand.
According to DeepSeek’s official announcement, off-peak pricing for DeepSeek-V4-Flash is set at 0.05 yuan for cached inputs, 1.5 yuan for uncached inputs, and 4.5 yuan for outputs per million tokens. During peak hours, all rates double, pushing V4-Pro’s cached input price to 0.3 yuan, uncached input to 9 yuan, and output to 27 yuan per million tokens. Compared to V4-Pro’s previous rate of 0.025 yuan (cached), 3 yuan (uncached), and 6 yuan (output), the peak-hour increase represents a maximum surge of 1,100%. This marks a departure from the company’s aggressive discounting strategy earlier this year. In late 2024, DeepSeek launched its V3 series, achieving performance comparable to leading models like GPT-4o at a fraction of the training cost. Its subsequent R1 model further disrupted the market, with output API pricing set at just 3% of OpenAI’s o1. By April, DeepSeek unveiled its V4 series—featuring the 1.6-trillion-parameter V4-Pro and the 284-billion-parameter V4-Flash—and slashed cached input prices to one-tenth of their launch value, with V4-Pro temporarily discounted to 0.025 yuan per million tokens, setting a global record for low pricing. V4-Flash’s cached input rate similarly dropped from 0.2 yuan to 0.02 yuan. Following a limited-time 2.5-fold discount that ended on May 31, V4-Pro’s official pricing was adjusted to one-quarter of its original list price, settling at 0.025 yuan (cached), 3 yuan (uncached), and 6 yuan (output) per million tokens.
DeepSeek’s ultra-low pricing rapidly drove massive adoption. Data from the global model distribution platform OpenRouter shows that between July 27 and August 2, DeepSeek-V4-Flash led global weekly API call volume with 7.22 trillion tokens, while V4-Pro ranked fourth. After V4-Flash’s official API entered public beta on July 31, it maintained the top spot for two consecutive weeks, logging 8.83 trillion tokens (August 3–9) and 11.2 trillion tokens (August 10–16). However, this explosive growth quickly strained infrastructure. On August 4, the open-source AI coding agent platform OpenCode reported that V4-Flash faced capacity shortages and potential error spikes due to unprecedented traffic, prompting emergency fixes. Reports from First Financial indicated that the official API was nearly unavailable that morning, though DeepSeek later confirmed the issue had been resolved. Behind the server pressure lies a broader compute shortage. At the 2026 China Development Forum Annual Conference, it was disclosed that China’s daily token call volume surpassed 140 trillion in March, representing a more than 1,000-fold increase from the 100 billion recorded in early 2024. Yet supply has failed to keep pace. High-end GPUs and HBM memory remain critically tight, with NVIDIA’s Blackwell and Hopper series facing severe supply constraints. Data from the China Academy of Information and Communications Technology (CAICT) reveals that domestic AI compute demand surged 417% year-over-year in Q1 2026, while effective supply growth lagged at just 128%, widening the gap. The strain previously forced Moonshot AI to pause new consumer subscriptions for Kimi K3, redirecting all available compute to existing subscribers.
DeepSeek’s pricing adjustment is not an isolated move but part of a broader industry trend. Throughout 2026, domestic large model providers have collectively raised prices. Zhipu AI, recognized as Hong Kong’s first listed LLM company, has implemented three price increases this year alone. Meanwhile, Moonshot AI significantly hiked API rates following the release of Kimi K3, with input prices rising over threefold and output prices nearing a fourfold increase. A Morgan Stanley research report published on August 9 noted that the average Chinese LLM API input price reached 4.9 yuan per million tokens in Q2 2026, with output prices at 21.9 yuan, reflecting increases of approximately 48% and 80% respectively compared to Q1 2025 figures of 3.3 yuan and 12.2 yuan. Despite the hikes, DeepSeek retains a significant cost advantage globally. Calculations by Guosheng Securities indicate that V4-Pro’s peak-hour output pricing remains roughly one-thirteenth that of Claude’s flagship model, dropping to approximately one-twenty-fifth during off-peak hours.
Critics who equate low pricing with compromised quality have found little evidence to support such claims. Official benchmarking data released by DeepSeek demonstrates that V4-Pro outperforms Opus-4.8 across multiple rigorous tests, including Terminal Bench 2.1, Cybergym, DeepSWE, and AutomationBench (Public). The model’s disruptive value proposition even spawned a new term on social media, as reported by the Jiangnan Metropolis Daily: the “DeepSeek kill line.” Derived from a cost-versus-intelligence scatter plot by Artificial Analysis, the chart uses V4-Flash as a dividing threshold. Models offering comparable performance at higher costs, or inferior performance at similar or higher prices, fall into a “kill zone,” signaling imminent obsolescence. Analysts argue that DeepSeek’s pivot away from relentless discounting signals a maturation of the domestic LLM sector, transitioning from price-driven competition to value-driven differentiation. Morgan Stanley emphasized that long-term competitiveness hinges on model intelligence rather than pricing. Sustained advancements require substantial investment in training and compute; relying solely on low prices compresses profit margins and starves future R&D cycles. Zheshang Securities echoed this sentiment, noting that DeepSeek’s leadership in normalizing pricing has shattered pessimistic expectations of an “infinite price war,” effectively establishing a new valuation anchor for China’s large model industry.