New Models
AI Minute Newsroom
2026-08-26
Show the robot one video of a ten-minute job, and it does the ten-minute job — no retraining, nothing typed in
Skild AI has released S1, a robot foundation model prompted with video instead of language: you show it a person doing a task and it produces the actions, in settings and on jobs it never saw during training. The company reports 66 percent success on unseen tasks at 100,000 hours of pre-training, against 9 percent for the same model prompted with words, and puts the value of one demonstration at roughly 380 episodes of post-training. It is the first such model shown holding together across ten-minute jobs with dozens of manipulation steps — making coffee, potting a plant, frying pancakes — and Skild says it improvises through mistakes rather than replaying the video. Performance still falls away when conditions shift far enough that the task needs a different strategy.
Why it mattersEvery robot deployment so far has been gated on collecting data for one specific task in one specific building. If a phone video is enough to specify the work, putting a robot on a new job stops being a data-collection project and becomes an afternoon. That is the step that decides whether these machines ever leave the demo reel.
✓ Verified · 2 sources
Read in the app — free, in 9 languages
Related stories
Microsoft's new model answers only with a number, never a sentence.
2026-10-10Shengshu's new video model charges 1.4 cents a second.
2026-10-09OpenAI's cheap model now has a mode costing six times more.
2026-10-09Perplexity's small model can search an index its big model built.
2026-10-09Anthropic's new small model costs 90% less than last year's.
2026-10-08