⚡ Uncle Cat AI Radar
IndustryModels

DeepSeek warns of a large across-the-board API price rise

The Chinese lab that set the global floor for inference pricing told developers it will raise API rates substantially, without naming figures or a date.

The notice

DeepSeek told developers through its open-platform console on Thursday that it plans to raise pricing across its API line "in the near term," and that the increase is expected to be large. The notice gave no new rate card and no effective date, saying only that formal details would follow. It advised companies and individual developers to plan their call volumes carefully and to avoid over-topping-up prepaid balances in the meantime.

The move reverses a two-year direction of travel. DeepSeek has cut prices repeatedly since 2024, most aggressively in May 2026, when it reduced rates by roughly three quarters. Current published pricing puts V4-Flash cache-hit input at 0.02 yuan per million tokens with output at 2 yuan, and V4-Pro at 0.025 yuan input and 6 yuan output — figures that made the company the reference point every other vendor was measured against. A peak-hour multiplier already doubles those rates on weekdays between 9:00–12:00 and 14:00–18:00 Beijing time.

What is driving it

The pressure is volume. V4-Flash, released under an MIT licence at the end of July, has become one of the most-called models anywhere; Chinese reports put single-day throughput on the platform at around 8 trillion tokens in early August, with cumulative call volume leading the domestic market. Serving that at near-cost rates consumes cash and, by the company's own account, has degraded stability and response quality during busy periods. DeepSeek is not alone: Zhipu raised prices sharply earlier this year, and several Chinese vendors have quietly retired their loss-leading tiers.

Why it matters

DeepSeek's pricing has functioned as a global anchor. Its cuts forced US labs to publish cheaper tiers, pushed open-weight hosting margins down across providers, and underwrote the economics of a generation of agent products that assume near-free tokens. A substantial increase from the price leader signals that the era of subsidised inference in China is ending, and that long-running agentic workloads — which burn tokens by the tens of millions per task — will be repriced accordingly. Because the weights remain open, the immediate effect may be to push heavy users toward self-hosting and third-party providers rather than to raise their bills outright.

Sources