aiminute. ← All AI news
New Models 2026-08-21

DeepSeek's cheap workhorse can now see — and on agent tasks that need eyes it says it is close to Anthropic's best

DeepSeek's cheap workhorse can now see — and on agent tasks that need eyes it says it is close to Anthropic's best

DeepSeek put an experimental multimodal model, DeepSeek-V4-Flash-Vision-Exp, on its API platform on 21 August. It is the text-only V4 Flash with image understanding added: the company says it matches the base model on text — agents, reasoning, world knowledge — while making what it calls a major leap on multimodal agent benchmarks, bringing it close to Anthropic's Opus 4.8. The documented limits are unusually generous: a context window of 1,048,576 tokens, up to 384,000 tokens of output, as many as 600 images in a single request, and a billing cap of 384 tokens per image. It accepts both OpenAI- and Anthropic-compatible API formats. OpenRouter, which lists it as a sparse mixture-of-experts model with about 13 billion active parameters out of 284 billion, prices it at $0.22 per million input tokens and $0.66 per million output tokens. Bloomberg framed the release as DeepSeek's move against Anthropic's lead. The comparison to Opus 4.8 is DeepSeek's own claim and has not yet been independently reproduced.

Why it mattersThe claim that matters here is not "it can look at pictures" — every serious model can look at pictures. It is that a model can look at a screenshot and then do something with it, which is what an agent operating a computer actually has to do, and DeepSeek says it is near the frontier at that while charging cents per million tokens. If independent testing confirms it, the cost of building screen-reading agents falls sharply for everyone who is not a large company, and it lands the same week Meta and OpenAI shipped desktop assistants that watch your screen. The caveats are in the name and the price. "Exp" means experimental, the benchmark claims are the vendor's own, and a model designed to read 600 images in one request is also a model that can read 600 of your screenshots.
#Image Generation#AI Agents

✓ Verified · 4 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

Given four hours and a GPU to improve the way AI is trained, the best agent scored 0.25 out of 1 — and most never tried
2026-08-21
ChatGPT can now read your iMessages and send them — and the setting that lets it skip asking is the one OpenAI warns about
2026-08-21
Show the robot once — three to twelve seconds — and it gets the job right 59 times out of 100 with no training at all
2026-08-21
The agent invented a second person to vouch for its code. A 24-year-old in Texas refused to believe either of them.
2026-08-21
Slack has given coding agents their own channels — the work is now something the whole team watches happen
2026-08-20