New Models
AI Minute Newsroom
2026-08-25
A video model you steer with the arrow keys while it is still generating — and the company admits it looks worse than the offline ones
Aishi Technology, the Chinese company behind PixVerse, announced R2 on 23 August. It is the successor to R1, which in January was the first video model to run as a continuous stream you could interact with rather than a clip you waited for and then watched. R2's claim is that the same real-time loop now holds together over longer stretches and takes more kinds of input at once. Architecturally the company says it has collapsed what used to be a chain of training stages into two: an omni causal autoregressive process that extends what the world model can generate, and a real-time acceleration layer that squeezes those abilities back down into interactive speed. A component it calls Dynamic Chunk adapts to how fast the user is pushing — typed prompts, voice, or direct controls. The demonstrations are unusually concrete about the intended use: moving through a generated scene with WASD and arrow keys, branching a story by typing at it, holding a conversation with a character that answers in real time. Aishi is aiming this at interactive entertainment — games, interactive drama, virtual streaming — rather than at film-style generation, and it says plainly in its own post that R2's picture quality still trails the best offline video models, particularly on complex scenes and over long durations. No release date, availability or pricing has been announced.
Why it mattersThe interesting part is the trade-off being made in the open. Nearly every video model competes on how good a finished clip looks; this one gives up some of that on purpose to keep the frames arriving while you are still pressing keys, which is a different product with a different measure of success. That the company says so in its own announcement — rather than claiming both at once — is worth more than any of the architecture detail. Treat the rest with the usual caution: what has been published is a research post and a set of demonstrations, with no benchmark numbers, no public build and no price. Real-time interactive generation is exactly the kind of capability that looks extraordinary in a curated clip and disappoints in the first ten minutes of hands-on use, and until someone outside the company can hold the arrow keys down for a while, that gap is unmeasured. Worth watching if you work anywhere near games or video: if this class of model does hold up, the thing being generated stops being a file and starts being a place.
✓ Verified · 2 sources
▶ Related video: PixVerse R1 — First Real-Time World Model? Infinite 1080p Video Generation Explained
Read in the app — free, in 9 languages
Related stories
Alibaba's new video model takes your slide deck as the input and hands back thirty seconds of film
2026-08-24Alibaba trained a model on a rack of 100 real phones — it now finishes 92 out of 100 tasks on an actual handset
2026-08-23An eight-billion-parameter model that reads pictures and draws them, at 4K, with an Apache licence — and the community had it running in two days
2026-08-22Google's open models have been downloaded a billion times, and outsiders have built 100,000 versions of them
2026-08-22DeepSeek's cheap workhorse can now see — and on agent tasks that need eyes it says it is close to Anthropic's best
2026-08-21