New Models
AI Minute Newsroom
2026-08-21
DeepSeek's cheap workhorse can now see — and on agent tasks that need eyes it says it is close to Anthropic's best
DeepSeek put an experimental multimodal model, DeepSeek-V4-Flash-Vision-Exp, on its API platform on 21 August. It is the text-only V4 Flash with image understanding added: the company says it matches the base model on text — agents, reasoning, world knowledge — while making what it calls a major leap on multimodal agent benchmarks, bringing it close to Anthropic's Opus 4.8. The documented limits are unusually generous: a context window of 1,048,576 tokens, up to 384,000 tokens of output, as many as 600 images in a single request, and a billing cap of 384 tokens per image. It accepts both OpenAI- and Anthropic-compatible API formats. OpenRouter, which lists it as a sparse mixture-of-experts model with about 13 billion active parameters out of 284 billion, prices it at $0.22 per million input tokens and $0.66 per million output tokens. Bloomberg framed the release as DeepSeek's move against Anthropic's lead. The comparison to Opus 4.8 is DeepSeek's own claim and has not yet been independently reproduced.
Why it mattersThe claim that matters here is not "it can look at pictures" — every serious model can look at pictures. It is that a model can look at a screenshot and then do something with it, which is what an agent operating a computer actually has to do, and DeepSeek says it is near the frontier at that while charging cents per million tokens. If independent testing confirms it, the cost of building screen-reading agents falls sharply for everyone who is not a large company, and it lands the same week Meta and OpenAI shipped desktop assistants that watch your screen. The caveats are in the name and the price. "Exp" means experimental, the benchmark claims are the vendor's own, and a model designed to read 600 images in one request is also a model that can read 600 of your screenshots.
✓ Verified · 4 sources
Read in the app — free, in 9 languages
Related stories
Microsoft's new transcriber starts writing before you finish speaking.
2026-10-05OpenAI will put ads beside the pictures you ask ChatGPT to draw.
2026-10-05An American lab is about to give away a China-class open model.
2026-10-05Meta's assistant keeps an hourly file on everyone in your life.
2026-10-05Two senators want prison time for bosses whose AI agents hack.
2026-10-05