DeepSeek V4 goes GA with 1M context as legacy APIs retire
DeepSeek's flagship V4 line exits preview: legacy chat and reasoner endpoints shut down today as new peak and off-peak API pricing takes effect.
DeepSeek moved its V4 model family to general availability on July 24, completing a transition announced at the end of June and hard-coded into its API: the legacy deepseek-chat and deepseek-reasoner endpoints stop responding today, and all traffic migrates to deepseek-v4-pro and deepseek-v4-flash.
From preview to production
The V4 line first appeared as an open-sourced preview on April 24 — V4-Pro at 1.6 trillion total parameters with 49 billion active, V4-Flash at 284 billion total with 13 billion active, both carrying a 1-million-token context window. According to TechNode, the general-availability build standardizes that million-token window across the lineup and targets stronger agentic task execution, mathematical reasoning, and code generation. Chinese media reported that a small-scale gray release began around July 21 after roughly two months of engineering optimization, with participating developers claiming overall quality approaching top closed-source models — claims that remain unverified pending independent benchmarks.
New pricing, new silicon
The GA release introduces DeepSeek's first peak/off-peak API pricing: calls during 9:00–12:00 and 14:00–18:00 Beijing time cost twice the off-peak rate, an unusual demand-shaping mechanism for an LLM API. Chinese outlets also report the official build has been adapted away from CUDA dependence and is compatible with multiple domestic AI chips from day one, consistent with earlier disclosures about V4 post-training runs on Huawei Ascend hardware. The API remains compatible with OpenAI and Anthropic SDKs, and DeepSeek documents integrations with coding agents including Claude Code, GitHub Copilot, and OpenCode.
Why it matters
DeepSeek remains the reference point for open-weight frontier models, and this release is less about new capability than about industrialization: a hard same-day migration, time-of-day pricing that treats inference like an electric grid, and domestic-chip portability all signal a lab now optimizing for sustained commercial scale rather than research splashes. For the crowded Chinese model market — where Kimi K3, Qwen3.8, and MiniMax M3 launched within days of each other — V4's GA resets the baseline that every rival's pricing and agent performance will be measured against.