aiminute. ← All AI news
Rumor Mill AI Minute Newsroom 2026-08-21

Rumour: the anonymous free model is still said to be Zhipu's — but the score that made it famous was ten tasks, and the full run is ordinary

Rumour: the anonymous free model is still said to be Zhipu's — but the score that made it famous was ten tasks, and the full run is ordinary

This is an unverified rumour, published as one. A model called Ox Alpha appeared on OpenRouter on 20 August under an anonymous provider labelled only 'Stealth' — free for a week, with a 1,048,576-token context window, up to 131,072 tokens of output, and text, image and video input. On 21 August the researcher Ben Davis published a fingerprint analysis saying he is '99% certain' the model is an unreleased member of Zhipu's GLM-5 family: video-encoder token consumption identical to GLM-5V-Turbo, a tokenizer matching GLM-5.3 apart from a fixed 75-token wrapper, the same refusal behaviour on audio input, the same '1301' error code, and a comparable emoji rate in its output. He says the analysis rules out Xiaomi MiMo, DeepSeek, Google, Qwen, xAI, OpenAI and Anthropic. None of it is confirmed: as of 23 August no company has claimed Ox Alpha and Zhipu has said nothing. Note also the data question. OpenRouter's own model page says prompts and completions sent to Ox Alpha are retained by the anonymous provider, though not used for training — while OpenCode promoted the same model as zero data retention. Until the provider is named, the safe assumption is that whatever you send is kept by a company that has not told you who it is. (Correction, 23 August: our original report said Ox Alpha scored 80% pass@1 on the DeepSWE coding benchmark, ahead of Claude Fable 5 at 65%, GLM-5.3 at 62% and GPT-5.6-sol at 52%. That figure came from a 10-task subset — under 9% of the benchmark. Davis has since run the full 113-task set and got roughly 63%, in line with GPT-5.6 Sol rather than ahead of the field, and he says the full-set number is the one that makes sense. Ox Alpha did not top the benchmark. Our headline said it did, and that was wrong.)

Why it mattersThe correction is the story. A 10-task result travelled around the world in two days as 'anonymous model tops coding benchmark'; the 113-task result that contradicts it will travel nowhere near as far. Whenever you see a model beat the field by a wide margin, the first question is how many problems it was actually asked — a small sample produces large, exciting, meaningless gaps. The rest still stands: the identity is unproven, and the pattern of launching a strong model anonymously and free for a week is a distribution channel, not a gift. You are the benchmark, and on this one the provider says it keeps what you send.
#Coding

⚠ Unverified rumor — treat with caution

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

OpenAI will ship a Codex upgrade daily for 28 days or reset limits.
2026-10-05
Google paused its open-source bug bounty after a flood of AI reports.
2026-10-04
Anthropic is paying to train 10,000 engineers to deploy Claude.
2026-10-04
A Linux desktop now refuses code written by AI outright.
2026-10-04
DeepSeek put its agent on your desktop and opened the code.
2026-10-03