On July 31, 2026, DeepSeek launched V4 Flash "0731" — a significant update to its budget model that climbed 10 points on the Artificial Analysis Intelligence Index, from 40 to 50 points, placing it just 1 point below OpenAI's GPT-5.6 Luna, released in July 2026.
On July 30, one day prior, OpenAI had cut the price of GPT-5.6 Luna by 80% and GPT-5.6 Terra by 20%. The temporal coincidence is not a coincidence — it is competitive pressure.
The math that the market needs to understand
DeepSeek V4 Flash-0731: $0.15/1M input tokens, $0.29/1M output tokens. GPT-5.6 Luna (after 80% cut): $0.20/1M input tokens, $1.20/1M output tokens. DeepSeek's output is approximately 4x cheaper.
In coding benchmarks, GPT-5.6 Luna delivers 67.2% on DeepSWE against 53.3% for V4 Flash — a 14 percentage point difference. The cost per task on DeepSWE: $0.61 for Luna, $0.10 for Flash. 6x cheaper.
The architecture of V4 Flash
284 billion total parameters, 13 billion active — a sparse MoE architecture with 256 routed experts and 6 active per token. Context window of 1 million tokens. MIT License — weights available on HuggingFace for anyone who wants to run locally. The "0731" update added native support for Responses-API and Codex, with a focus on agents and coding.
What this means for the market
DeepSeek went back to doing what it did in January 2026 with V3: forcing the market to rethink the cost-quality ratio. The difference now is that proprietary models have already cut prices in response — proving that the pressure works.
For teams with cost-sensitive workloads: V4 Flash. For high-precision workloads integrated into the OpenAI ecosystem: GPT-5.6 Luna. For most agents in production: a combination of both — Flash as a low-cost filter, Luna for high-demand tasks.

