DeepSeek image recognition mode is officially launched: a key step towards the productization of multi-modal capabilities

DeepSeek's image recognition mode is officially launched on the web and App, using the Thinking with Visual Primitives framework, alongside "Quick Mode" and "Expert Mode", marking the transition of multi-modal capabilities from research and development to productization.

DeepSeek multi-modal researcher Xiaokang Chen announced that DeepSeek’s image recognition mode has been officially launched on the web and App. "Image Reading Mode" is juxtaposed with "Quick Mode" and "Expert Mode". When turned on, users can directly upload images for DeepSeek to understand visual content. Its capabilities exceed simple text extraction (OCR).

DeepSeek image recognition mode

The underlying technology framework for the launch of this image recognition mode is the "Thinking with Visual Primitives" core framework released by DeepSeek in April this year. The core idea of ​​the framework is to allow the model to decompose visual information into "primitives" for reasoning like humans do, rather than simply mapping images to text end-to-end. This technical route is unique among domestic multi-modal models.

Visual primitive framework

The launch of the image recognition mode is a key step for DeepSeek to move from "king of pure text reasoning" to "multi-modal all-round player". Previously, DeepSeek's core competitiveness focused on mathematical reasoning, code generation and long text processing, and its visual capabilities have always been its obvious shortcomings in competition with competing products such as Doubao.

There was also a dramatic episode on the day it was launched - according to The Paper, DeepSeek's image recognition mode was unable to correctly identify photos of founder Liang Wenfeng, but instead identified him as Dong Yuhui, Zhang Xuefeng and even Lei Jun. Lei Jun's photo was accurately identified. One explanation is that Liang Wenfeng acts extremely low-key, and there are few public photos and information on the Internet, making it difficult for the model to form stable identification features. This "not knowing the boss" oolong just proves that DeepSeek has not done special recognition optimization for its own boss - the model's face recognition ability is based on the distribution of public training data, not internal privileges.

The official release of this image recognition mode, combined with the overall upgrade of the V4 series models, completes DeepSeek’s multi-modal puzzle. What is worthy of attention in the future is whether the performance of visual understanding in actual scenes can reach the same level as the reasoning capabilities of the V4 series, and whether it will overlap with the multi-modal version planned in V4.1.

Copyright: Content sourced from IT Home . This platform has compiled and organized this content for informational purposes and learning exchange only. If there are any copyright concerns, please contact us for resolution.

Reviews

  • Loading reviews...