Cloudflare published the playbook it uses to serve heavyweight open models like Moonshot's Kimi K2.6 and Z.ai's GLM 5.2 on its Workers AI platform. Storing the attention cache in 8-bit instead of 16-bit doubles context capacity to 1.37 million tokens and delivers 41% higher throughput at roughly 30% lower cost per token; compressing GLM's weights to 4-bit shrinks the model from 705 GB to 421 GB. Accuracy on benchmark suites was indistinguishable from the originals, the company says.