Cloudflare doubles context for Kimi K2.6
Cloudflare cut serving costs for Kimi K2.6 and GLM 5.2 with KV cache quantization and INT4 weight compression, roughly doubling context and cutting size 40%.
Cloudflare detailed three inference optimizations for serving Moonshot’s Kimi and Z.ai’s GLM model families on its Workers AI platform, published August 3. FP8 KV-cache quantization roughly doubled usable context on Kimi K2.6, from about 686,000 to 1.37 million tokens, while lifting peak throughput 41% at 64 concurrent requests, with negligible accuracy loss. INT4 weight compression cut GLM 5.2’s checkpoint size 40% (705GB to 421GB) with decode-throughput gains from 16% to 55% depending on concurrency, and a new KV-cache integrity check guards against corruption across concurrent requests for under 1% overhead. The stack runs on the open-source SGLang framework across disaggregated prefill/decode pools on H200 GPUs.