aiminute. ← All AI news
New Models AI Minute Newsroom 2026-08-23

Alibaba trained a model on a rack of 100 real phones — it now finishes 92 out of 100 tasks on an actual handset

Alibaba trained a model on a rack of 100 real phones — it now finishes 92 out of 100 tasks on an actual handset

Alibaba's Tongyi MAI team released Qwen-UI-Agent on 20 August, a model built to operate screens rather than answer questions — phones, desktops, browsers, and a command line, in one system. Instead of training only in simulation, the team ran a live rig of more than 100 physical smartphones covering 150-plus apps, and built MobileWorld-Real, a benchmark of 400-plus tasks on real devices. Reported results: 92.2 percent on MobileWorld-Real, 97.5 percent on AndroidDaily, 79.5 percent on OSWorld-Verified and 73.6 percent on WebArena, ahead of Claude Opus 4.8 and GPT-5.6 on mobile and browser work — though Opus still leads on the harder OSWorld-v2, 54.8 to 40. The smaller MAI-UI models at 2B and 8B are on Hugging Face under Apache 2.0.

Why it mattersThe gap between a demo agent and a useful one has always been the sim-to-real gap: an agent that clicks correctly in a simulator gets lost on a real phone with a real popup. Training against a rack of actual handsets is an expensive, unglamorous way to close it, and the numbers on real devices are the ones worth watching. The small versions being open means the technique does not stay inside Alibaba.
#Coding#AI Agents

✓ Verified · 3 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

An AI designed an open chip that runs small language models.
2026-10-07
A video model now labels footage shot from a robot's own view.
2026-10-07
OpenAI's new endpoint picks from a fixed list in 150 milliseconds.
2026-10-07
Google's new search model fits in 191 megabytes of phone memory.
2026-10-07
Developers think a free mystery model belongs to Thinking Machines.
2026-10-06