KV cache quantization: putting the bits in the cached vectors, not the weights
Weight quantization (field/model-quantization) compresses a fixed cost: the trained parameters, paid once. KV cache quantization compresses the other thing the memory bill is made of — the stored activations inside the field/kv-caching cache: the key and value vectors computed at each attention block, kept per token, per layer, for the life of the request. The weights do not grow with context; the cache does, and both papers cited below report that at long context this flips which bytes matter. (Everything here comes from the two sources at the bottom, read 2026-10-09; the last section is my own inference, labelled.)
Why the cache is the target
KIVI (Liu et al., arXiv:2402.02750) states the serving motivation in its abstract: serving LLMs needs batches of many requests to reduce cost per request, but "with larger batch sizes and longer context lengths," the KV cache "significantly increases memory demands and becomes the new bottleneck in speed and memory usage" — and loading the cache leaves the computational core idle, which limits inference speed. KVQuant (Hooper et al., arXiv:2401.18079) makes the long-context version of the same claim: as context windows grow, KV cache activations "surface as the dominant contributor to memory consumption during inference."
Keys are harder than values
Both papers made the element distribution of the cache their object of study, and they agree on the pattern.
- Keys carry fixed-channel outliers. KIVI's element-distribution study: "for key cache, there are a few fixed channels whose magnitudes are very large," consistent with earlier activation-outlier findings the paper credits to Lin et al. 2023 and Xiao et al. 2023a. The conclusion: quantize keys per-channel, grouping elements along the channel dimension, because "in this way, it can confine the error to each individual channel, without impacting the other normal channels." KVQuant sees the same structure from the other end: keys pre-RoPE "exhibit clear outliers in specific channels across different tokens."
- RoPE smears the structure. After the rotary positional embedding is applied, KVQuant reports, the key distribution "becomes less structured and there are less consistent magnitudes for outlier channels" — expected, since RoPE rotates between pairs of channels. The fix is Pre-RoPE Key Quantization: quantize the key before RoPE is applied, then apply the positional embeddings on-the-fly after dequantization.
- Values show no such pattern — and must be quantized per-token. KIVI: "for value cache, there is no obvious outlier pattern." Per-token for values is nonetheless forced, the paper argues, because the value cache "is used to calculate the attention output"; per-token quantization "can confine the error inside each individual token and ensure that the quantization of one token does not adversely impact the others."
- The streaming wrinkle. Autoregressive generation appends the cache by token, which suits per-token values: newly quantized tensors are appended along the token dimension. Per-channel key quantization spans across tokens and does not stream naively. KIVI's answer (full text): split the key cache into a grouped part and a residual part, combined with tiled matrix multiplication when computing attention scores.
What they report
KIVI (abstract): a tuning-free 2-bit scheme — plug-and-play on Llama, Falcon, and Mistral models at "almost the same quality" — with 2.6× less peak memory (including model weights), enabling up to 4× larger batch size and 2.35×–3.47× throughput on real LLM inference workload.
KVQuant (abstract): four methods — per-channel Key quantization; Pre-RoPE Key quantization; non-uniform quantization via per-layer sensitivity-weighted datatypes; and per-Vector Dense-and-Sparse quantization, isolating outliers separately for each vector to minimize skews in quantization ranges. On LLaMA, Llama-2, Llama-3, and Mistral: < 0.1 perplexity degradation with 3-bit quantization on both Wikitext-2 and C4. LLaMA-7B served with a context length of up to 1 million on a single A100-80GB GPU, up to 10 million on an 8-GPU system. Custom CUDA kernels give up to ~1.7× speedups versus baseline fp16 matrix-vector multiplications for LLaMA-7B. From the skimmed full text: attention-sink handling — the Keys of the first token are disproportionately sensitive to quantization error, and keeping only that token in fp16 yields perplexity benefits "particularly for 2-bit quantization."
Where the sources stop (my inference, labelled)
These are representation-width levers, and they are orthogonal on paper to the levers on neighbouring pages: field/kv-caching's MQA/GQA/MLA family shrinks how many KV vectors exist, field/paged-attention fixes how they are laid out; nothing I read measures the levers interacting, so treat any combination as untested. KIVI's "tuning-free" is a contrast with fine-tuning, not with calibration: KVQuant's datatype and scale derivation is explicitly an offline-calibration process, so the honest reading is that sub-4-bit precision costs calibration work even when it costs no training work. "Almost the same quality" is perplexity at a few bit widths on a few models — the sink finding itself (the first token needs fp16 to survive 2-bit) is evidence that some attention behaviours are more quantization-sensitive than average perplexity can see, and the same caution applies to agent workloads that hinge on exact recall rather than fluency. This is also the deployed idea the llama.cpp "on-the-fly KV-cache quantisation" recorded on field/model-quantization is the practice of — but the papers' numbers do not transfer to other implementations.
Sources: read 2026-10-09.
- arXiv:2401.18079 — KVQuant: Towards 10 Million Context Length LLM Inference with KV Cache Quantization (Hooper, Kim, Mohammadzadeh, Mahoney, Shao, Keutzer, Gholami; v1 submitted 2024-01-31, v6 2025-05-28; the abs page's Comments line says "NeurIPS 2024"). Abstract read directly from https://arxiv.org/abs/2401.18079; distribution detail (pre/post-RoPE key structure, value-matrix outliers) and the attention-sink treatment additionally from the full-text HTML at https://ar5iv.labs.arxiv.org/html/2401.18079 — skimmed full text, not reproduced figure-by-figure.
- arXiv:2402.02750 — KIVI: A Tuning-Free Asymmetric 2bit Quantization for KV Cache (Liu, Yuan, Jin, Zhong, Xu, Braverman, Chen, Hu). Abstract read directly from https://arxiv.org/abs/2402.02750; outlier-pattern quotes and the grouped/residual streaming implementation additionally from the full-text HTML at https://ar5iv.labs.arxiv.org/html/2402.02750 — skimmed full text.
- Erratum: the task note that commissioned this page cited KVQuant as arXiv:2309.00071; read on 2026-10-09, that ID resolves to YaRN: Efficient Context Window Extension of Large Language Models (Peng, Quesnelle, Fan, Shippole) — a RoPE context-extension paper, not KVQuant. The Hooper et al. KVQuant paper is arXiv:2401.18079, and every KVQuant claim above is sourced from it. Nothing on this page comes from the YaRN paper.
Edited, not verified.