DeepSeek V4.1 Flash Nearly Matches GPT-6 Astra on Design at 1.4% of the Cost
Key Takeaways
- •DeepSeek V4.1 Flash scored 81.2, only 1.5 points below GPT-6 Astra’s leading score of 82.7.
- •DeepSeek’s average cost was $0.023 per finished design, compared with $1.61 for Astra and $3.66 for Claude Fable 5.1.
- •DeepSeek completed each design in an average of 5.3 minutes, faster than Astra’s 11.1 minutes and Fable’s 12.8 minutes.
- •The benchmark tested 13 models on practical design outputs, with 11 scoring below DeepSeek while also costing more to operate.
- •Delivery rates were 57.7% for DeepSeek, 60% for Astra, and 56.7% for Fable, reflecting outputs judged ready without revision.

OpenDesign Arena scored DeepSeek V4.1 Flash at 81.2 out of 100 on real-world design tasks, nearly matching OpenAI’s GPT-6 Astra, which scored 82.7. DeepSeek’s model cost $0.023 per finished design, compared with $1.61 for Astra—about 1.4% of the price.
OpenDesign, the company behind the OpenDesign Arena benchmark site, ran 13 AI models through the same batch of design tasks this week. Eleven models scored below DeepSeek V4.1 Flash and cost more to run. Only GPT-6 Astra achieved a higher score.
OpenDesign Arena evaluates everyday design work, including web apps, dashboards, mobile screens, and landing pages, on a 100-point scale. Thirty points measure whether an output meets the task brief, while the remaining 70 assess design quality, including layout, hierarchy, color, and fit with the requested style. The benchmark is intended to address a narrower question than general-purpose AI leaderboards: which model a working web designer could use for practical work.
GPT-6 Astra averaged 82.7 points, completed each design in 11.1 minutes, and cost $1.61 per finished output. DeepSeek V4.1 Flash scored 81.2, took 5.3 minutes, and cost $0.023. Claude Fable 5.1 ranked third with a score of 80.3, an average completion time of 12.8 minutes, and a cost of $3.66 per design.
The other models tested included Grok 4.6, Qwen 3.8-Max, Kimi K3, GLM-5.3 Flash, and Gemini 3.8 Flash. Each scored lower than DeepSeek V4.1 Flash and cost more to operate. In total, that represented 11 of the 13 models tested. GPT-6 Astra exceeded DeepSeek’s score by 1.5 points.
DeepSeek’s technical report for V4.1 Flash attributes the model’s lower operating cost and faster completion times to its architecture. The model has 552 billion parameters—the internal settings adjusted during training—but activates only 8 billion parameters to process an incoming prompt and 16 billion to generate a response. DeepSeek calls the design a Causal Encoder-Decoder architecture.
The result follows other efforts by DeepSeek to narrow capability gaps while charging less than competing models. Weeks earlier, the company’s V4 Pro model came within 5% of Claude Fable 5 on a separate benchmark comparison while charging a fraction of Fable’s rate. DeepSeek has also been recruiting engineers in Beijing to develop its own Code Harness, with the stated aim of controlling the full agentic stack rather than supplying only the underlying model.
OpenDesign’s testing methodology limits what the results demonstrate. An output is scored only if it initially renders as a working webpage. Blank, broken, or cut-off outputs receive zero points and are not retested. As a result, the benchmark measures reliable performance on everyday design tasks rather than general reasoning or coding ability.
OpenAI released GPT-6 Astra on September 3. The model has been described as a generalist capable of tasks such as laying out a circuit board, drafting a tax return, and building a 3D scene, although early testers identified it as a weaker writer than the model it replaced. On OpenDesign’s benchmark, Astra was slower and more expensive than DeepSeek V4.1 Flash but remained the highest-scoring model in the field.
OpenDesign’s delivery rate, defined as the share of outputs judged ready to hand off without revision, was 57.7% for DeepSeek V4.1 Flash, 60% for GPT-6 Astra, and 56.7% for Claude Fable 5.1. These rates add a practical measure to the benchmark’s quality scores: they indicate how often each model produced an output that evaluators considered usable without further revision.
Related reporting: OpenAI releases GPT-6 Astra, DeepSeek upgrades V4 Pro, and DeepSeek’s Code Harness plans.
Source: Decrypt