NewsStocksDeepSeek Unveils Experimental Vision Model, Citing Internal Tests Near Anthropic's Opus 4.8

DeepSeek Unveils Experimental Vision Model, Citing Internal Tests Near Anthropic's Opus 4.8

Author: Coincentral·

Key Takeaways

  • DeepSeek announced DeepSeek-V4-Flash-Vision-Exp on August 21, 2026, adding image and screenshot processing to its V4 Flash model family for the first time.
  • The experimental model retains the text, reasoning, agentic and general knowledge capabilities of DeepSeek-V4-Flash and is available to developers through DeepSeek's API.
  • DeepSeek's own internal evaluations showed performance similar to Anthropic's Opus 4.8 in multimodal agent tests, but third-party benchmarks are still needed to verify the claims.
  • The launch occurs amid escalating AI competition between China and the United States, including U.S. scrutiny of Chinese developers and China's national AI investment plan through 2030.
  • Anthropic is reportedly preparing for a possible initial public offering that could be filed as early as this month, with Morgan Stanley, Goldman Sachs and JPMorgan said to be involved in preparations.
DeepSeek Unveils Experimental Vision Model, Citing Internal Tests Near Anthropic's Opus 4.8

China's DeepSeek has introduced an experimental multimodal AI model that can process images and screenshots, expanding its V4 model family as competition among the world's major artificial intelligence developers — in both China and the United States — continues to intensify.

The release, announced on August 21, 2026, adds visual processing to the V4 line for the first time, while the existing text system remains in place rather than being replaced. Image input has become a standard capability across frontier models from OpenAI, Google and Anthropic, and the addition brings DeepSeek's newest model family in line with that baseline.

DeepSeek Adds Vision Capabilities to V4 Flash

DeepSeek announced DeepSeek-V4-Flash-Vision-Exp, an experimental model that adds visual processing to its V4 Flash platform. Users can upload images and screenshots, then ask the model to analyze the content or complete tasks based on what it sees.

The company said the new model keeps the text capabilities of DeepSeek-V4-Flash, including reasoning, agentic tasks and general knowledge. DeepSeek described the release as a multimodal extension of its newer model family rather than a replacement for its existing text system.

"This experimental multimodal model matches DeepSeek-V4-Flash on text capabilities—including agents, reasoning, and world knowledge," DeepSeek said.

The model is available through DeepSeek's application programming interface, allowing developers to add image-based functions — such as analyzing uploaded screenshots — to their own products and services and to build visual AI applications on top of the platform.

News of the release was shared on X by Walter Bloomberg (@DeItaone):

DEEPSEEK CHALLENGES ANTHROPIC WITH NEW AI MODEL

DeepSeek has unveiled an experimental multimodal AI model capable of analyzing images, screenshots and text while performing autonomous tasks.

Called DeepSeek-V4-Flash-Vision-Exp, the model reportedly approaches Anthropic's Opus…

Walter Bloomberg (@DeItaone) August 21, 2026

DeepSeek Compares Model With Anthropic Opus 4.8

DeepSeek said internal evaluations showed its new model performing at levels similar to Anthropic's Opus 4.8 in multimodal agent tests. These tests measure how AI systems handle tasks that require both visual input and limited human guidance. The comparison places the experimental model directly against one of Anthropic's higher-end systems.

Anthropic has emphasized agentic capabilities across its Claude line, including a computer-use feature introduced in late 2024 that allows the model to interpret what appears on a screen and carry out tasks inside applications — overlapping with the screenshot-driven tasks DeepSeek's new model is built to perform.

However, DeepSeek's claims are based on its own testing, and broader third-party benchmarks will be needed to compare performance across different workloads.

DeepSeek has already developed vision-language models through its DeepSeek-VL family, an earlier line of models that combine image understanding with language processing. The latest release brings similar visual functions into the newer V4 line, expanding the range of tasks the platform can handle. The company, an AI lab backed by the Chinese quantitative hedge fund High-Flyer, first drew global attention in January 2025 when its R1 reasoning model delivered results competitive with leading U.S. systems at a fraction of their reported training costs.

AI Competition Grows Across China and the U.S.

The launch comes as Chinese AI companies continue to expand model development across text, vision and agent capabilities. China has also introduced a national plan through 2030 that calls for greater investment in AI chips, large computing systems, multimodal models, autonomous agents and blockchain technology.

Chinese developers are also facing growing scrutiny from U.S. officials over competition in advanced AI. Recent debate has included concerns about intellectual property, model training methods and access to the high-end computing hardware needed to train large models.

Meanwhile, Anthropic is reportedly preparing for a possible initial public offering. The company could file as early as this month, with Morgan Stanley, Goldman Sachs and JPMorgan said to be involved in preparations.

The potential listing has drawn attention after SpaceX completed an IPO that raised about $86.2 billion, including its overallotment. Comparisons between the two companies remain speculative because Anthropic has not announced final offering terms or a confirmed valuation.

DeepSeek's new model enters that competitive environment as developers in China and the United States continue expanding multimodal AI systems — models that combine text, images and autonomous task execution — for both consumer and enterprise use.