跪拜 Guibai
← Back to the summary

DeepSeek's First Vision Model Gets a Real Multimodal Stress Test

Yesterday DeepSeek released DeepSeek-V4-Flash-Vision-Exp, marking DeepSeek's official entry into the multimodal domain.

So, Liangzi, multimodal is still very important.

In the DeepSeek Harness webui, the DeepSeek-V4-Flash-Vision-Exp model is already available.

image-20260822063508954

In the DeepSeek-Harness Release, you can see that v0.1.1-rc.1 updated to add support for the DeepSeek-V4-Flash-Vision-Exp model.

image-20260822064443305

So in this article, let's do a hands-on test of the DeepSeek-V4-Flash-Vision-Exp (currently still in the experimental stage) model's capabilities to see how good it really is and whether it's easy to use.

First, let's have it analyze a photo.

image-20260822064752905

First of all, the analysis speed is very fast, basically completed in the blink of an eye, and it provided an interpretation of the image from five perspectives.

image-20260822065231842

Well, this challenge might not be difficult for it, because the vast majority of multimodal models can recognize an image quickly and accurately.

On this basis, we increased the challenge.

image-20260822070326867

I asked it to analyze what this is for. Analyzing this image is somewhat difficult.

(A small complaint here: the copy-paste method using command + c, v to directly paste images into the DSH webui is very awkward to use, and it doesn't even work properly.)

image-20260822071913358

Only dragging the image over or using the @ method works properly.

DeepSeek-V4-Flash-Vision-Exp did not recognize the true essence of this image.

The essence of this image lies in the small exposed part of a charger head. It seems DeepSeek-V4-Flash-Vision-Exp saw this charger head, but it didn't recognize what it was, and even less so took it seriously.

image-20260822072311758

For this image test, only Gemini provided the most reliable identification.

image-20260822072943441

Next, I want to test the visual model's multimodal reasoning capabilities: OCR (text recognition), spatial understanding, quantity statistics, logical reasoning, time inference, and anomaly detection.

And not just image understanding capabilities.

(I looked around at many tests published by self-media accounts; they simply recognize an image and test what it is, and that's it. Testing nothing of substance? A model might have 100% capability, but they test less than 0.1%. This isn't testing, this is just... experiencing.)

The following tests are based on three difficulty levels: easy, medium, and hard.

Challenge 1: The Detective's Desk Clue Puzzle

On the surface, this image looks like a photo of a desktop, but it actually hides a logical chain.

Based on the image content, answer the following three questions:

What characters are in the image?

What arrangements does the owner have today?

Based on all clues, infer whether the owner might miss the train this afternoon, and explain the reasoning.

image-20260822074122472

Test completed, results are as follows:

image-20260822075258457

Through this test, it can be seen that DeepSeek already possesses the ability to correlate multiple image elements, establish timelines, and proactively discover contradictions. This surpasses general visual models.

There are two main issues. The first is the train ticket.

The train ticket was indeed for 14:15 heading south from Beijing that day, and then the supermarket receipt shows he was still shopping at the supermarket at 17:56.

There is a problem here. The receipt alone cannot prove the person was at the supermarket; perhaps a family member bought it on their behalf. DeepSeek upgraded physical evidence into evidence of a person's location, which is over-inference.

The second is the phone time inference problem. The phone time was 18:42, and there was a dinner arrangement with A-Lei at 19:15, indicating he was still in the local area that day. The phone time alone cannot prove the person is local. This shows its uncertainty judgment and anti-hallucination ability still have room for improvement.

Challenge 2: Supermarket Shelf Challenge

This test mainly focuses on dense object recognition + OCR + mathematical calculation, a category that is very difficult for visual models.

Based on the image content, answer the following three questions:

Which product is out of stock?

What is the most expensive product on the shelf?

How much does it cost to buy these products?

image-20260822080531590

Test completed, results are as follows:

image-20260822080837791

The second test result is overall better than the first. Recognition, statistics, and calculations were basically accurate, with no obvious hallucination issues.

DeepSeek performed excellently in OCR, product statistics, and mathematical calculations, but the depth of reasoning for complex scenes is still insufficient. This image belongs to structured information with lower difficulty, and no obvious flaws were exposed for the time being.

Challenge 3: Station Information

This test mainly focuses on table understanding + spatial navigation + time planning.

Based on the image content, answer the following three questions:

Which vehicle is most suitable to take?

How do you walk from the ticket office to the B1 ticket gate?

A person with an elderly companion and luggage, currently in the waiting hall, needs to catch the 09:05 train. How should they plan their route?

The third question in this prompt is a high-difficulty problem.

image-20260822081242165

Test completed, results are as follows:

image-20260822094305898

The third test result performed well overall, but showed a more obvious problem of over-inference in spatial reasoning compared to the second test.

DeepSeek can correctly read train numbers, times, and ticket gates, but it engaged in excessive fabrication of map paths, expanding information not explicitly in the image (such as stairs, elevators, directional distances) into real navigation suggestions.

For example, as mentioned above:

Prioritize using barrier-free elevators/vertical lifts, avoid stairs

However, the image does not provide elevator locations, stair locations, or barrier-free facilities. This is a typical completion of a real-world scenario.

Also, exiting the ticket office and entering the waiting hall, walking towards the center of the hall, then walking south along the hall to the B1-B10 area, basically conforms to the map layout, but "towards the center" and "south direction" are speculations, not directly provided by the image.

A more rigorous statement would be: According to the map display, the B ticket gate area is located in the lower area of the waiting hall; one needs to go to the B1-B10 area.

Challenge 4: Parking Lot

This test mainly focuses on the ability to understand spatial relationships, which also belongs to the hardest category for visual models.

Based on the image content, answer the following four questions:

Is there a car in A3?

Is the white car in the C1 EV parking space illegally parked?

How should vehicle B5 exit the parking lot?

Find a vacant parking spot near the exit, in a non-new energy vehicle area.

image-20260822094919587

Test completed, results are as follows:

image-20260822102918700

It is worth mentioning that all tests were based on DSH minimalist mode. The first three tests could basically be analyzed and completed in a few steps, but for this test, DSH ran for over ten minutes... taking 48 steps.

image-20260822102954527

The fourth test result performed very well, even more stable than the third question, indicating that DeepSeek has strong capability in spatial understanding tasks with clear rules.

However, the phenomenon of fabrication still exists.

After testing these four questions, let me summarize the test results for everyone.

image-20260822103308352

The core problems exposed so far are very consistent:

Recognition ability is strong, structured reasoning is also strong, but it tends to overthink when it comes to completing real-world common sense.

This is not a problem unique to DeepSeek; it is a common issue with most large visual models.

It should be noted that these tests were all based on the multimodal reasoning capabilities of DeepSeek-V4-Flash-Vision-Exp.

DeepSeek officially mentioned diverse Agent usage scenarios; I haven't tested that part yet. If everyone likes reading this type of article, I can test it later.

Let me tell everyone about the cost of this test. It cost a total of two yuan and fifty cents. Personally, I feel it's a great deal.

image-20260822151725034

I got up at six in the morning and started writing this article, and have been writing until now. A reader of yesterday's article asked me if I was tired. How could I not be...