aiminute. ← All AI news
Tools 2026-08-04

Cloudflare shows how it squeezes trillion-parameter open models onto fewer GPUs

Cloudflare shows how it squeezes trillion-parameter open models onto fewer GPUs

Cloudflare published the playbook it uses to serve heavyweight open models like Moonshot's Kimi K2.6 and Z.ai's GLM 5.2 on its Workers AI platform. Storing the attention cache in 8-bit instead of 16-bit doubles context capacity to 1.37 million tokens and delivers 41% higher throughput at roughly 30% lower cost per token; compressing GLM's weights to 4-bit shrinks the model from 705 GB to 421 GB. Accuracy on benchmark suites was indistinguishable from the originals, the company says.

Why it mattersThe frontier of AI economics is not only better models — it is serving the same models cheaper. Techniques like these are a big reason heavyweight open models keep getting cheaper to run, and Cloudflare published them openly for any infrastructure team to copy.

✓ Verified · 1 sources

WhatsApp X Telegram
Read in the app — free, in 9 languages

Related stories

Apple Music will tell you when a song was made by a machine — but the uploader decides whether to say so
2026-08-21
Stripe has just paid $7.5 billion for a model router. Days later Ramp built one and is giving it away until January.
2026-08-21
Meta's assistant is now a Mac app that reads your screen and types into any window — and what it sees can train the model
2026-08-21
One click on a news site now tells Google to show you more of it — and you stay on the page you were reading
2026-08-21
Adobe will now generate the music, the voiceover and the door slam — and it says the licence covers you
2026-08-21