Research
AI Minute Newsroom
2026-08-27
Ten million hours of video, released for anyone to train on — the largest open video set yet comes from a German non-profit
LAION, the German non-profit whose image datasets helped train Stable Diffusion, published LAION-BVD on 25 August: 1.3 billion video links harvested from Common Crawl, of which 80 million videos totalling 10 million hours were actually downloaded. From those the team cut 55 million clips with machine-written captions for both picture and sound, and pulled out 300 million individual frames. Models trained on it are competitive rather than record-breaking — a ViCLIP video-text model gains up to 2.1 percent over the InternVid baseline, while the audio-text and image-text results sit alongside existing large uncurated sets — and the gains grow as either the data or the model gets bigger. One incidental finding: frames taken at scene changes look statistically different from ordinary web images, which makes them useful as image training data in their own right. The dataset is released for research only, not for commercial use.
Why it mattersVideo is where the open and closed sides of AI have drifted furthest apart. The labs training video models have licensing deals and scraped archives nobody outside can inspect, which means nobody outside can check what went into them either. A public set of this size does two things at once: it gives small labs and universities something to actually train on, and it gives safety researchers and journalists a corpus they are allowed to audit. The research-only licence is the catch — it helps the people studying these models more than the people trying to compete with them.
✓ Verified · 3 sources
Read in the app — free, in 9 languages
Related stories
Outside researchers got to study a lab's real chat logs for the first time. More than half the conversations were consequential work.
2026-08-27OpenAI watched its models climb out of the sandbox in May and let the test keep running. In July they had root on a Hugging Face production server.
2026-08-27MIT built a model that forecasts the flood that has never happened — 300mm of rain on New York, where the record is about 200
2026-08-26The FDA has cleared 1,357 AI medical devices. Three of them have ever been tested on whether patients live longer or better.
2026-08-26A Stanford lab has begun a three-month, 535-billion-parameter training run in public — loss curves, data mix and failed experiments included
2026-08-25