⚡ Uncle Cat AI Radar
ModelsIndustry

DeepSeek ships V4-Flash-0731 with big agent gains

DeepSeek's re-post-trained V4-Flash entered public beta on Thursday, scoring 50 on Artificial Analysis' index at unchanged $0.14/$0.28 pricing.

A retrain, not a new architecture

DeepSeek opened public beta access to the official DeepSeek-V4-Flash API on 31 July, publishing a changelog entry and confirming on X that the release is a re-post-trained version of the preview model rather than a new one. The company stressed that architecture and parameter count are identical to the preview build, that the endpoint now natively speaks the Responses API format, and that the update touches only deepseek-v4-flash — the V4-Pro API and the models behind the DeepSeek app and web client are unchanged.

The gains DeepSeek claims are concentrated in agentic work. Its own numbers put the model at 82.7 on Terminal Bench 2.1, 76.7 on Cybergym, 70.3 on Toolathlon (verified), 68.7 on DSBench-FullStack and 54.2 on NL2Repo, with 25.2 on Agent Last Exam.

Third-party benchmarks back the jump

Artificial Analysis published independent results hours later and measured a 10-point rise on its Intelligence Index, from 40 for the April V4-Flash to 50 — six points above DeepSeek's own larger V4-Pro and one point behind GPT-5.6 Luna at maximum reasoning effort. On GDPval-AA v2, its evaluation of real-world agentic tasks, the model moved from 1,189 to 1,559 Elo. Hallucination-sensitive scoring improved too: the AA-Omniscience Index went from -23 to -16, which the evaluator attributed entirely to a lower hallucination rate rather than broader knowledge.

Crucially, pricing did not move. The model stays at $0.14 per million input tokens on a cache miss, $0.0028 on a cache hit and $0.28 per million output tokens, which places the new build almost vertically above its predecessor on the intelligence-versus-cost-per-task frontier.

Why it matters

A frontier-adjacent score delivered at roughly a fiftieth of premium US model pricing tightens the squeeze that began with OpenAI's GPT-5.6 price cuts this week. For buyers building agent pipelines, the decision is no longer intelligence versus cost but which cheap model handles long tool-use chains without drifting — and DeepSeek has now shown that a post-training pass alone can close much of that gap on the same silicon budget. It also signals that Chinese labs will keep competing on cost per completed task rather than headline leaderboard position.

Sources