Study: KV cache compression beats more GPUs
Archive item — written before sources were shown.
New research shows compressing the KV cache is 1.2 to 2x cheaper than tensor parallelism across configurations for serving large-context models.
A new arXiv paper compares two ways to serve large language models under long-context, memory-constrained conditions: adding more GPUs via tensor parallelism, or compressing the key-value cache to fit on fewer. Across the configurations tested, the researchers found KV cache compression is 1.2 to 2 times cheaper than scaling out with tensor parallelism, and the paper identifies model-size thresholds where one approach clearly beats the other, giving operators a concrete decision rule instead of a default toward more hardware.
- 01More GPUs or Smaller Cache? Tensor Parallelism vs. KV Compressionarxiv.org · primary
