New Models
AI Minute Newsroom
2026-08-01
MiniMax's H3 makes 2K video with built-in stereo sound — at a third of rivals' price
Shanghai-based MiniMax released H3 on July 31, a video model that reads text, images, video and audio as one combined input and returns clips of 5 to 15 seconds at up to 2K and 24fps — with voice, sound effects and music generated together as native stereo audio rather than added afterwards. It can also edit existing footage and transfer motion between videos. MiniMax says 2K generation costs less than a third of mainstream rivals, and that the model weights will be opened 'in the coming days'.
Why it mattersVideo generation is racing toward the same fate as text: better, cheaper, open. Sound produced jointly with the picture removes one of AI video's last telltale seams, and if the promised open weights arrive, the toolkit for both creators and fakers gets dramatically cheaper — pricing pressure on Sora, Veo and Runway included.
✓ Verified · 3 sources
Read in the app — free, in 9 languages
Related stories
Reflection will hand out a 501-billion-parameter model for free.
2026-10-06Microsoft's new transcriber starts writing before you finish speaking.
2026-10-05ElevenLabs is giving students a free year of its reading voice.
2026-10-05An American lab is about to give away a China-class open model.
2026-10-05A video site now gives away the best-scoring open translation model.
2026-10-05