aiminute. ← All AI news
New Models AI Minute Newsroom 2026-08-25

A video model you steer with the arrow keys while it is still generating — and the company admits it looks worse than the offline ones

A video model you steer with the arrow keys while it is still generating — and the company admits it looks worse than the offline ones

Aishi Technology, the Chinese company behind PixVerse, announced R2 on 23 August. It is the successor to R1, which in January was the first video model to run as a continuous stream you could interact with rather than a clip you waited for and then watched. R2's claim is that the same real-time loop now holds together over longer stretches and takes more kinds of input at once. Architecturally the company says it has collapsed what used to be a chain of training stages into two: an omni causal autoregressive process that extends what the world model can generate, and a real-time acceleration layer that squeezes those abilities back down into interactive speed. A component it calls Dynamic Chunk adapts to how fast the user is pushing — typed prompts, voice, or direct controls. The demonstrations are unusually concrete about the intended use: moving through a generated scene with WASD and arrow keys, branching a story by typing at it, holding a conversation with a character that answers in real time. Aishi is aiming this at interactive entertainment — games, interactive drama, virtual streaming — rather than at film-style generation, and it says plainly in its own post that R2's picture quality still trails the best offline video models, particularly on complex scenes and over long durations. No release date, availability or pricing has been announced.

Why it mattersThe interesting part is the trade-off being made in the open. Nearly every video model competes on how good a finished clip looks; this one gives up some of that on purpose to keep the frames arriving while you are still pressing keys, which is a different product with a different measure of success. That the company says so in its own announcement — rather than claiming both at once — is worth more than any of the architecture detail. Treat the rest with the usual caution: what has been published is a research post and a set of demonstrations, with no benchmark numbers, no public build and no price. Real-time interactive generation is exactly the kind of capability that looks extraordinary in a curated clip and disappoints in the first ten minutes of hands-on use, and until someone outside the company can hold the arrow keys down for a while, that gap is unmeasured. Worth watching if you work anywhere near games or video: if this class of model does hold up, the thing being generated stops being a file and starts being a place.
#Video

✓ Verified · 2 sources

▶ Related video: PixVerse R1 — First Real-Time World Model? Infinite 1080p Video Generation Explained
WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

OpenAI's cheap model now has a mode costing six times more.
2026-10-09
Perplexity's small model can search an index its big model built.
2026-10-09
Anthropic's new small model costs 90% less than last year's.
2026-10-08
A video model now labels footage shot from a robot's own view.
2026-10-07
Google's new search model fits in 191 megabytes of phone memory.
2026-10-07