Google TurboQuant: 3-Bit KV Cache Compression Explained

Google TurboQuant compresses the LLM KV cache to about 3 bits per value with no measured accuracy loss and cuts KV memory by at least 6x. How PolarQuant and QJL work, the paper's benchmarks, the vLLM --kv-cache-dtype turboquant presets, llama.cpp and SGLang status, and Qdrant vector search.

October 7, 2026 · 1 min read

TL;DR

Last updated October 7, 2026.

6x
“TurboQuant reduces key-value memory by at least 6x on needle-in-a-haystack tests while keeping perfect downstream results.”
Google Research blog, March 24, 2026

Google TurboQuant is a quantization algorithm that shrinks the KV cache of an LLM to about 3 bits per value with no training or calibration step. The paper (arXiv 2504.19874, an ICLR 2026 paper) reports quality neutrality at 3.5 bits per channel and marginal degradation at 2.5 bits. Google's announcement measured at least a 6x smaller KV cache and up to 8x faster attention-logit computation for 4-bit keys on H100 GPUs.

What it is

A data-oblivious vector quantizer. It rotates each vector, applies a fixed per-coordinate quantizer, and spends 1 extra bit on a QJL sketch of the error. It works on KV caches and on vector-search embeddings.

Where you can run it

vLLM has a TurboQuant attention backend (merged April 15, 2026). Qdrant 1.18 ships it for vector search. llama.cpp mainline has the rotation stage only.

What Is Google TurboQuant?

TurboQuant comes from Amir Zandieh and Vahab Mirrokni's group at Google Research. The arXiv paper, TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate, was first posted April 28, 2025 by Zandieh, Majid Daliri, Majid Hadian, and Mirrokni. Google's blog post on March 24, 2026 brought it to a wide audience, alongside the two building blocks it uses: QJL (June 2024) and PolarQuant (February 2025, an AISTATS 2026 paper).

The problem it targets is memory overhead in ordinary quantization. Block-wise schemes store a scale and zero point in full precision for every small block of numbers, and Google estimates that bookkeeping adds 1 or 2 bits per number. At a 3-bit budget that overhead eats a large share of the savings. TurboQuant removes the per-block constants entirely.

The paper also proves a lower bound on the distortion any vector quantizer can reach, and shows TurboQuant lands within a constant factor of about 2.7 of that bound at every bit width. That guarantee holds for any input, because the method never looks at the data distribution before quantizing.

3.5 bits
Quality-neutral KV cache
2.5 bits
Marginal degradation
6x+
KV memory reduction (NIAH)
8x
Attention logits, 4-bit on H100

How TurboQuant Works: PolarQuant and QJL

Stage one is a random rotation. Multiplying a vector by a random orthogonal matrix spreads its energy evenly, and in high dimensions each coordinate of the rotated vector follows a concentrated Beta distribution whose shape is known ahead of time. Because the distribution is fixed, TurboQuant precomputes an optimal scalar quantizer (a codebook) for each bit width and applies it to every coordinate independently. Nothing data-dependent gets stored next to the codes.

Google's blog describes this stage through PolarQuant, which converts pairs of coordinates into a radius and an angle. After rotation the angles cluster in a predictable range, so the grid boundaries are known before any data arrives and the normalization step that block quantizers need disappears.

Stage two fixes a subtle problem. A quantizer tuned for mean squared error gives biased estimates of dot products, and attention scores are dot products between queries and cached keys. TurboQuant spends 1 more bit per coordinate on a Quantized Johnson-Lindenstrauss transform of the residual: it projects the leftover error and keeps only the sign of each projected value. QJL pairs that 1-bit sketch with a full-precision query, which makes the inner-product estimate unbiased.

Online and data-oblivious

TurboQuant does not need a calibration set or fine-tuning, and it never inspects the model. Each key and value is quantized the moment it is written to the cache, including during token-by-token generation. Serving engines can therefore adopt it as a drop-in KV cache format.

TurboQuant Benchmarks

Google evaluated TurboQuant, PolarQuant, and QJL on LongBench, Needle In A Haystack, ZeroSCROLLS, RULER, and L-Eval using open models from the Gemma, Mistral, and Llama families. The needle test from the paper is the clearest single comparison. Every method ran at a memory compression ratio of 0.25, meaning it kept a quarter of the full KV cache.

Needle-In-A-Haystack score, Llama-3.1-8B-Instruct, 25% of full KV cache memory (TurboQuant paper, Figure 4)
MethodTypeScore
Full precisionUncompressed baseline0.997
TurboQuantVector quantization0.997
PolarQuantVector quantization0.995
KIVIScalar quantization0.981
PyramidKVToken eviction0.895
SnapKVToken eviction0.858

TurboQuant matched the uncompressed cache exactly. Token-eviction methods (SnapKV, PyramidKV) lost more because they drop whole tokens, and a needle sitting in a dropped token cannot be recovered. On LongBench the paper ran TurboQuant at 2.5 and 3.5 bits during streaming generation, splitting channels into outlier and normal sets and giving outliers more bits, which is where the fractional bit widths come from.

For speed, Google measured attention-logit computation against an optimized JAX baseline: 4-bit TurboQuant keys ran up to 8x faster than 32-bit unquantized keys on H100 accelerators. Smaller keys mean less memory traffic, and decode-time attention is bound by memory bandwidth.

TurboQuant KV Cache Memory Savings

KV cache size per token is 2 × layers × KV heads × head dimension × bytes per value. Llama-3.1-8B has 32 layers, 8 KV heads, and a head dimension of 128, so it stores 65,536 numbers per token. The table below applies that to a full 128K-token context. It counts the quantized values only and leaves out small metadata such as vector norms. Our KV cache guide walks through the formula for other models.

KV cache for one 128K-token sequence, Llama-3.1-8B (computed from the model config)
FormatBits per valuePer token128K tokens
BF16 (default)16128 KiB16 GiB
FP8 KV cache864 KiB8 GiB
4-bit432 KiB4 GiB
TurboQuant 3.5-bit3.528 KiB3.5 GiB
TurboQuant 2.5-bit2.520 KiB2.5 GiB

Going from BF16 to 3.5 bits frees about 12.5 GiB per long sequence on an 8B model. On a server that memory turns into more concurrent requests or longer contexts on the same GPU, which is where the cost savings show up.

TurboQuant in vLLM

vLLM added TurboQuant as an attention backend in pull request 38479, merged April 15, 2026. It quantizes at store time with fused Triton kernels and needs no offline calibration or change to model weights. The implementation follows the paper's rotation idea with a Walsh-Hadamard transform and random sign flips, uses Lloyd-Max scalar quantization for keys, and stores values with uniform quantization. Pull request 39931 (merged May 5, 2026) added hybrid models and uniform quantization.

Serve a model with a TurboQuant KV cache in vLLM

vllm serve Qwen/Qwen3-4B --kv-cache-dtype turboquant_k8v4

# keep boundary layers at full precision
vllm serve Qwen/Qwen3-4B --kv-cache-dtype turboquant_k8v4 \
  --kv-cache-dtype-skip-layers 0,1,34,35
vLLM TurboQuant presets, Qwen3-4B, head_dim 128 (PR 38479). Baseline: GSM8K 0.900, NIAH 100%
PresetKeysValuesCompressionGSM8KNIAH
turboquant_k8v4FP8 (E4M3)4-bit uniform2.6x0.860100%
turboquant_4bit_nc4-bit MSE + NC4-bit uniform + NC3.8x0.840100%
turboquant_k3v4_nc3-bit MSE + NC4-bit uniform + NC4.3x0.780100%
turboquant_3bit_nc3-bit MSE + NC3-bit uniform + NC4.9x0.720100%

The PR's throughput tests on four RTX PRO 6000 Blackwell GPUs put turboquant_k8v4 at 79% to 100% of baseline output tokens per second and turboquant_4bit_nc at 71% to 96%. Decode-heavy requests lose the most, landing at 65% to 79% of baseline across presets. Long prefill runs at parity: the 8,192-in, 64-out scenario hit 100% of baseline with k8v4. The trade is capacity for per-request speed, so TurboQuant suits workloads limited by KV memory, such as long agent contexts and high concurrency.

Accuracy depends on the preset

Retrieval held at 100% on every preset, but GSM8K reasoning fell from 0.900 to 0.720 at 3 bits on Qwen3-4B. Run your own evals on the model and preset you plan to deploy before switching a production cache.

TurboQuant in llama.cpp and SGLang

llama.cpp mainline does not support TurboQuant as of October 7, 2026. The feature request opened March 25, 2026 was closed as not planned. Mainline did merge pull request 21038 on April 1, 2026, which adds a Hadamard rotation to the KV cache before quantization and cites TurboQuant. The TurboQuant+ project describes rotation plus the stock q4_0 cache as TurboQuant's rotation stage. The PolarQuant codebook and the rest of the method still live in forks.

SGLang received four TurboQuant KV cache pull requests between March and April 2026, including a fused Triton version reporting 3.88x compression. All four were closed without being merged.

TurboQuant GitHub Implementations

Google published the method through the paper and blog. Most code people use comes from independent repositories that appeared within days of the March announcement. Star counts below were checked October 7, 2026.

Popular TurboQuant repositories on GitHub
RepositoryWhat it isStars
RyanCodrai/turbovecVector index built on TurboQuant, Rust with Python bindings17,351
TheTom/turboquant_plusReference implementation and benchmarks, 3.8x to 6.4x KV cache compression7,028
0xSero/turboquantKV cache quantization with a Triton kernel (3-bit keys, 2-bit values)1,791
tonbistudio/turboquant-pytorchFrom-scratch PyTorch implementation for KV cache compression1,047
mitkox/vllm-turboquantvLLM integration fork618

TurboQuant vs FP8 KV Cache and KIVI

FP8 KV cache vs TurboQuant
FP8 KV cacheTurboQuant
Bits per value82.5 to 4
Compression vs BF162x2.6x to 6x+
Stored scalesPer tensor, or per head with calibrationNone
CalibrationRecommended in vLLM docsNot needed
vLLM flag--kv-cache-dtype fp8--kv-cache-dtype turboquant_k8v4 (and 3 more presets)

FP8 is the safe default: half the memory and the widest kernel support. TurboQuant is the option when 2x is not enough and you need 4x or more. Against KIVI, a widely used 2-bit KV cache quantizer, the paper's needle test at a quarter of full memory scored 0.997 for TurboQuant and 0.981 for KIVI on Llama-3.1-8B-Instruct.

Pros and Cons

Strengths
  • About 3 bits per value with no measured accuracy loss in Google's long-context tests
  • No calibration data, fine-tuning, or model changes
  • Provable distortion bound within a constant factor of about 2.7 of optimal
  • Upstream vLLM backend with four presets
  • Same method works for vector search (Qdrant 1.18)
Limitations
  • Dequantization costs throughput on short, decode-heavy requests (65% to 80% of baseline in vLLM tests)
  • Aggressive presets lose reasoning accuracy on small models
  • llama.cpp mainline has only the rotation stage, and SGLang PRs were closed unmerged
  • Most implementations people run are community projects, not Google code

FAQ

What is Google TurboQuant?

TurboQuant is a vector quantization algorithm from Google Research, announced March 24, 2026 and published as an ICLR 2026 paper (arXiv 2504.19874). It compresses LLM key-value cache vectors and vector-search embeddings to a few bits per number without calibration data or retraining. The paper reports quality neutrality on KV cache quantization at 3.5 bits per channel and marginal degradation at 2.5 bits per channel.

How does TurboQuant work?

TurboQuant randomly rotates each input vector, which gives every coordinate a concentrated Beta distribution that is known in advance. A precomputed optimal scalar quantizer is then applied to each coordinate, so no per-block scale or zero point has to be stored. A second stage applies a 1-bit Quantized Johnson-Lindenstrauss (QJL) transform to the residual error, which makes the inner-product estimate used by attention unbiased.

How much memory does TurboQuant save?

Google reports that TurboQuant reduces KV cache memory by at least 6x on needle-in-a-haystack tests while keeping perfect retrieval scores. At 3.5 bits per value, the KV cache for 128K tokens of Llama-3.1-8B drops from about 16 GiB in BF16 to about 3.5 GiB. vLLM's production presets are more conservative and report 2.6x to 4.9x compression on Qwen3-4B.

Does vLLM support TurboQuant?

Yes. vLLM merged a TurboQuant attention backend in pull request 38479 on April 15, 2026. Enable it with --kv-cache-dtype turboquant_k8v4 (FP8 keys and 4-bit values, about 2.6x compression) or the more aggressive turboquant_4bit_nc, turboquant_k3v4_nc, and turboquant_3bit_nc presets. Pull request 39931, merged May 5, 2026, added hybrid-model support.

Does llama.cpp support TurboQuant?

Partly. The llama.cpp feature request for TurboQuant (issue 20977) was closed as not planned. Mainline did merge a Hadamard rotation of the KV cache before quantization (pull request 21038, April 1, 2026), which cites TurboQuant and covers its rotation stage when combined with the stock q4_0 cache. The PolarQuant codebook and the full method remain in community forks such as the TurboQuant+ project.

Is TurboQuant lossless?

TurboQuant is a lossy quantizer, but its error is small enough that Google measured no accuracy loss at about 3 bits on long-context benchmarks. On needle-in-a-haystack with Llama-3.1-8B-Instruct, TurboQuant scored 0.997, the same as the full-precision cache. Lower bit widths and real serving stacks can lose quality: vLLM's 3-bit preset scored 0.720 on GSM8K versus a 0.900 baseline on Qwen3-4B.

Related Articles

Private deployments

The fastest endpoints are private deployments

Morph's top speeds come from dedicated deployments, not shared public endpoints: speculators trained on your traffic, caching tuned to your workload, and volume discounts over public per-token rates. Over 500 billion tokens per day run this way.

Talk to us about a private deployment

Serve long-context models without managing the KV cache

Morph serves open models such as GLM-5.3 and DeepSeek V4.1 Flash on custom kernels behind one OpenAI-compatible API.

Sources