
While OpenAI are busy celebrating the reduced price for GPT-5.6 Luna, the perennial disruptors at DeepSeek are quietly launching a beta of their much cheaper and highly anticipated V4 model — the DeepSeek V4 Flash 0731.
The model lands precisely one point behind Luna on the Artificial Analysis benchmark, and is a «a significant step up from the previous generation,» the DeepSeek V4 Flash (40), AA writes.
It has a one million token context window, and has jumped from forty to fifty points on the AA evaluation and is now on par with Gemini 3.6 Flash, just behind GLM-5.2, Muse Spark, and GPT Luna.
The biggest selling point is the price, which is at $0.14/$0.28 for one million tokens in and out, far below anything on the market for this kind of performance/cost.
Compared with Luna’s newly reduced prices of $0.30/$1.20, it ticks in at about half the cost for input tokens and around a fourth of the cost for output.
The model also has a cache discount of 98%, costing just $0.0028 for cache hits.
DeepSeek V4 Flash 0731 is now available on Hugging Face with weights under an MIT License and in the API.
Read More: Hugging Face page, Artificial Analysis rundown, and The Decoder. Discussion on r/Singularity and Hacker News.












