DeepSeek's new V4.1 Flash model delivers performance close to OpenAI's GPT-6 Astra in an evaluation of practical design work while sharply lowering costs, the results showed.
On Sept. 10 (local time), OpenDesign Arena gave DeepSeek V4.1 Flash a score of 81.2 out of 100, Decrypt and other outlets reported. That was 1.5 points below GPT-6 Astra's 82.7, or about 98 percent of its level.
The cost gap was larger. GPT-6 Astra cost $1.61 per finished design, while DeepSeek V4.1 Flash cost $0.023. That is about 1.4 percent of GPT-6 Astra on a price basis. It was also faster. GPT-6 Astra took an average of 11.1 minutes, while V4.1 Flash completed the task in 5.3 minutes.
The comparison applied the same design task to 13 AI models, OpenDesign said. The evaluated models included Anthropic's Claude Fable 5.1, Grok 4.6 and Qwen 3.8-Max. Only GPT-6 Astra scored higher than V4.1 Flash. The other 11 models scored lower than V4.1 Flash and cost more.
OpenDesign Arena evaluates models based on routine design work. It grades web apps, dashboards, mobile screens and landing page production tasks out of 100 points, allocating 30 points for meeting requirements and 70 points for layout and information hierarchy, color and fit with style. The evaluation focused on the ability to reliably produce outputs in real design work rather than general reasoning or coding skills.
Claude Fable 5.1 scored 80.3. Its average processing time was 12.8 minutes and its cost per finished output was $3.66. V4.1 Flash showed an advantage in speed and cost without a wide score gap.
DeepSeek presented a new model structure as the backdrop to such cost efficiency. According to a technical document, V4.1 Flash is a mixture-of-experts (MoE) model with a total of 552 billion parameters. Of those, 8 billion parameters are activated when processing input, and 16 billion when generating output. DeepSeek described it as a "causal encoder-decoder" structure. The key is boosting cost efficiency by designing different scales of activated parameters for input and output.
Still, there are limits to what the evaluation proves. In OpenDesign's test, outputs are graded only if they are rendered as a functioning webpage. If a blank screen appears, a page breaks, or the output is cut off midway, it is scored zero and not retested. As a result, the benchmark measures how stably a model can produce reliable, routine design outputs rather than general-purpose reasoning or coding ability.
GPT-6 Astra, which OpenAI unveiled on Sept. 3, targets a range of areas including computer use, software engineering, cybersecurity, science and professional work. It recorded the top score in the design benchmark as well, but lagged behind DeepSeek V4.1 Flash in speed and cost.
The gap was also not large in the share judged ready to hand over immediately after completion without separate revisions. V4.1 Flash recorded 57.7 percent, GPT-6 Astra 60 percent, and Claude Fable 5.1 56.7 percent.
DeepSeek has recently continued a strategy of narrowing the performance gap with rival models by emphasizing low costs. V4 Pro was assessed in a separate benchmark as showing a gap of around 5 percent with Claude Fable 5, and its price was also lower. However, that comparison combined multiple benchmarks, and results can differ depending on evaluation methods.
DeepSeek also recruited product managers and R&D engineers in Beijing in May to develop Code Harness. It is seen as a move to expand beyond the model itself into code agents and an execution environment for developers.
Introducing DeepSeek-V4.1-Flash: smarter, faster, more efficient. Introducing the smallest model in our new architecture family, with native visual understanding. Designed for greater capability, faster inference, higher throughput, and scaling to larger models. 1/6 pic.twitter.com/wxJGiyX56o