TL;DR
Last updated October 7, 2026.
“TurboQuant reduces key-value memory by at least 6x on needle-in-a-haystack tests while keeping perfect downstream results.”
Google TurboQuant is a quantization algorithm that shrinks the KV cache of an LLM to about 3 bits per value with no training or calibration step. The paper (arXiv 2504.19874, an ICLR 2026 paper) reports quality neutrality at 3.5 bits per channel and marginal degradation at 2.5 bits. Google's announcement measured at least a 6x smaller KV cache and up to 8x faster attention-logit computation for 4-bit keys on H100 GPUs.
What it is
A data-oblivious vector quantizer. It rotates each vector, applies a fixed per-coordinate quantizer, and spends 1 extra bit on a QJL sketch of the error. It works on KV caches and on vector-search embeddings.
Where you can run it
vLLM has a TurboQuant attention backend (merged April 15, 2026). Qdrant 1.18 ships it for vector search. llama.cpp mainline has the rotation stage only.
What Is Google TurboQuant?
TurboQuant comes from Amir Zandieh and Vahab Mirrokni's group at Google Research. The arXiv paper, TurboQuant: Online Vector Quantization with Near-optimal Distortion Rate, was first posted April 28, 2025 by Zandieh, Majid Daliri, Majid Hadian, and Mirrokni. Google's blog post on March 24, 2026 brought it to a wide audience, alongside the two building blocks it uses: QJL (June 2024) and PolarQuant (February 2025, an AISTATS 2026 paper).
The problem it targets is memory overhead in ordinary quantization. Block-wise schemes store a scale and zero point in full precision for every small block of numbers, and Google estimates that bookkeeping adds 1 or 2 bits per number. At a 3-bit budget that overhead eats a large share of the savings. TurboQuant removes the per-block constants entirely.
The paper also proves a lower bound on the distortion any vector quantizer can reach, and shows TurboQuant lands within a constant factor of about 2.7 of that bound at every bit width. That guarantee holds for any input, because the method never looks at the data distribution before quantizing.
How TurboQuant Works: PolarQuant and QJL
Stage one is a random rotation. Multiplying a vector by a random orthogonal matrix spreads its energy evenly, and in high dimensions each coordinate of the rotated vector follows a concentrated Beta distribution whose shape is known ahead of time. Because the distribution is fixed, TurboQuant precomputes an optimal scalar quantizer (a codebook) for each bit width and applies it to every coordinate independently. Nothing data-dependent gets stored next to the codes.
Google's blog describes this stage through PolarQuant, which converts pairs of coordinates into a radius and an angle. After rotation the angles cluster in a predictable range, so the grid boundaries are known before any data arrives and the normalization step that block quantizers need disappears.
Stage two fixes a subtle problem. A quantizer tuned for mean squared error gives biased estimates of dot products, and attention scores are dot products between queries and cached keys. TurboQuant spends 1 more bit per coordinate on a Quantized Johnson-Lindenstrauss transform of the residual: it projects the leftover error and keeps only the sign of each projected value. QJL pairs that 1-bit sketch with a full-precision query, which makes the inner-product estimate unbiased.
TurboQuant does not need a calibration set or fine-tuning, and it never inspects the model. Each key and value is quantized the moment it is written to the cache, including during token-by-token generation. Serving engines can therefore adopt it as a drop-in KV cache format.
TurboQuant Benchmarks
Google evaluated TurboQuant, PolarQuant, and QJL on LongBench, Needle In A Haystack, ZeroSCROLLS, RULER, and L-Eval using open models from the Gemma, Mistral, and Llama families. The needle test from the paper is the clearest single comparison. Every method ran at a memory compression ratio of 0.25, meaning it kept a quarter of the full KV cache.
| Method | Type | Score |
|---|---|---|
| Full precision | Uncompressed baseline | 0.997 |
| TurboQuant | Vector quantization | 0.997 |
| PolarQuant | Vector quantization | 0.995 |
| KIVI | Scalar quantization | 0.981 |
| PyramidKV | Token eviction | 0.895 |
| SnapKV | Token eviction | 0.858 |
TurboQuant matched the uncompressed cache exactly. Token-eviction methods (SnapKV, PyramidKV) lost more because they drop whole tokens, and a needle sitting in a dropped token cannot be recovered. On LongBench the paper ran TurboQuant at 2.5 and 3.5 bits during streaming generation, splitting channels into outlier and normal sets and giving outliers more bits, which is where the fractional bit widths come from.
For speed, Google measured attention-logit computation against an optimized JAX baseline: 4-bit TurboQuant keys ran up to 8x faster than 32-bit unquantized keys on H100 accelerators. Smaller keys mean less memory traffic, and decode-time attention is bound by memory bandwidth.
TurboQuant KV Cache Memory Savings
KV cache size per token is 2 × layers × KV heads × head dimension × bytes per value. Llama-3.1-8B has 32 layers, 8 KV heads, and a head dimension of 128, so it stores 65,536 numbers per token. The table below applies that to a full 128K-token context. It counts the quantized values only and leaves out small metadata such as vector norms. Our KV cache guide walks through the formula for other models.
| Format | Bits per value | Per token | 128K tokens |
|---|---|---|---|
| BF16 (default) | 16 | 128 KiB | 16 GiB |
| FP8 KV cache | 8 | 64 KiB | 8 GiB |
| 4-bit | 4 | 32 KiB | 4 GiB |
| TurboQuant 3.5-bit | 3.5 | 28 KiB | 3.5 GiB |
| TurboQuant 2.5-bit | 2.5 | 20 KiB | 2.5 GiB |
Going from BF16 to 3.5 bits frees about 12.5 GiB per long sequence on an 8B model. On a server that memory turns into more concurrent requests or longer contexts on the same GPU, which is where the cost savings show up.
TurboQuant in vLLM
vLLM added TurboQuant as an attention backend in pull request 38479, merged April 15, 2026. It quantizes at store time with fused Triton kernels and needs no offline calibration or change to model weights. The implementation follows the paper's rotation idea with a Walsh-Hadamard transform and random sign flips, uses Lloyd-Max scalar quantization for keys, and stores values with uniform quantization. Pull request 39931 (merged May 5, 2026) added hybrid models and uniform quantization.
Serve a model with a TurboQuant KV cache in vLLM
vllm serve Qwen/Qwen3-4B --kv-cache-dtype turboquant_k8v4
# keep boundary layers at full precision
vllm serve Qwen/Qwen3-4B --kv-cache-dtype turboquant_k8v4 \
--kv-cache-dtype-skip-layers 0,1,34,35| Preset | Keys | Values | Compression | GSM8K | NIAH |
|---|---|---|---|---|---|
| turboquant_k8v4 | FP8 (E4M3) | 4-bit uniform | 2.6x | 0.860 | 100% |
| turboquant_4bit_nc | 4-bit MSE + NC | 4-bit uniform + NC | 3.8x | 0.840 | 100% |
| turboquant_k3v4_nc | 3-bit MSE + NC | 4-bit uniform + NC | 4.3x | 0.780 | 100% |
| turboquant_3bit_nc | 3-bit MSE + NC | 3-bit uniform + NC | 4.9x | 0.720 | 100% |
The PR's throughput tests on four RTX PRO 6000 Blackwell GPUs put turboquant_k8v4 at 79% to 100% of baseline output tokens per second and turboquant_4bit_nc at 71% to 96%. Decode-heavy requests lose the most, landing at 65% to 79% of baseline across presets. Long prefill runs at parity: the 8,192-in, 64-out scenario hit 100% of baseline with k8v4. The trade is capacity for per-request speed, so TurboQuant suits workloads limited by KV memory, such as long agent contexts and high concurrency.
Retrieval held at 100% on every preset, but GSM8K reasoning fell from 0.900 to 0.720 at 3 bits on Qwen3-4B. Run your own evals on the model and preset you plan to deploy before switching a production cache.
TurboQuant in llama.cpp and SGLang
llama.cpp mainline does not support TurboQuant as of October 7, 2026. The feature request opened March 25, 2026 was closed as not planned. Mainline did merge pull request 21038 on April 1, 2026, which adds a Hadamard rotation to the KV cache before quantization and cites TurboQuant. The TurboQuant+ project describes rotation plus the stock q4_0 cache as TurboQuant's rotation stage. The PolarQuant codebook and the rest of the method still live in forks.
SGLang received four TurboQuant KV cache pull requests between March and April 2026, including a fused Triton version reporting 3.88x compression. All four were closed without being merged.
TurboQuant GitHub Implementations
Google published the method through the paper and blog. Most code people use comes from independent repositories that appeared within days of the March announcement. Star counts below were checked October 7, 2026.
| Repository | What it is | Stars |
|---|---|---|
| RyanCodrai/turbovec | Vector index built on TurboQuant, Rust with Python bindings | 17,351 |
| TheTom/turboquant_plus | Reference implementation and benchmarks, 3.8x to 6.4x KV cache compression | 7,028 |
| 0xSero/turboquant | KV cache quantization with a Triton kernel (3-bit keys, 2-bit values) | 1,791 |
| tonbistudio/turboquant-pytorch | From-scratch PyTorch implementation for KV cache compression | 1,047 |
| mitkox/vllm-turboquant | vLLM integration fork | 618 |
TurboQuant for Vector Search
The same quantizer compresses embeddings in a vector database. In the paper, TurboQuant beat the PQ and RabbiQ baselines named in Google's post on 1@k recall on GloVe (d=200), while indexing time dropped to nearly zero because there is no codebook to train on the dataset.
Qdrant 1.18 ships TurboQuant as a production quantization option. Qdrant's May 13, 2026 benchmarks found 4-bit TurboQuant within about 1 to 2 percentage points of its scalar quantization (which compresses 4x), and 2-bit and 1-bit TurboQuant giving higher recall than binary quantization at the same storage budget.
TurboQuant vs FP8 KV Cache and KIVI
| FP8 KV cache | TurboQuant | |
|---|---|---|
| Bits per value | 8 | 2.5 to 4 |
| Compression vs BF16 | 2x | 2.6x to 6x+ |
| Stored scales | Per tensor, or per head with calibration | None |
| Calibration | Recommended in vLLM docs | Not needed |
| vLLM flag | --kv-cache-dtype fp8 | --kv-cache-dtype turboquant_k8v4 (and 3 more presets) |
FP8 is the safe default: half the memory and the widest kernel support. TurboQuant is the option when 2x is not enough and you need 4x or more. Against KIVI, a widely used 2-bit KV cache quantizer, the paper's needle test at a quarter of full memory scored 0.997 for TurboQuant and 0.981 for KIVI on Llama-3.1-8B-Instruct.
Pros and Cons
- About 3 bits per value with no measured accuracy loss in Google's long-context tests
- No calibration data, fine-tuning, or model changes
- Provable distortion bound within a constant factor of about 2.7 of optimal
- Upstream vLLM backend with four presets
- Same method works for vector search (Qdrant 1.18)
- Dequantization costs throughput on short, decode-heavy requests (65% to 80% of baseline in vLLM tests)
- Aggressive presets lose reasoning accuracy on small models
- llama.cpp mainline has only the rotation stage, and SGLang PRs were closed unmerged
- Most implementations people run are community projects, not Google code
FAQ
What is Google TurboQuant?
TurboQuant is a vector quantization algorithm from Google Research, announced March 24, 2026 and published as an ICLR 2026 paper (arXiv 2504.19874). It compresses LLM key-value cache vectors and vector-search embeddings to a few bits per number without calibration data or retraining. The paper reports quality neutrality on KV cache quantization at 3.5 bits per channel and marginal degradation at 2.5 bits per channel.
How does TurboQuant work?
TurboQuant randomly rotates each input vector, which gives every coordinate a concentrated Beta distribution that is known in advance. A precomputed optimal scalar quantizer is then applied to each coordinate, so no per-block scale or zero point has to be stored. A second stage applies a 1-bit Quantized Johnson-Lindenstrauss (QJL) transform to the residual error, which makes the inner-product estimate used by attention unbiased.
How much memory does TurboQuant save?
Google reports that TurboQuant reduces KV cache memory by at least 6x on needle-in-a-haystack tests while keeping perfect retrieval scores. At 3.5 bits per value, the KV cache for 128K tokens of Llama-3.1-8B drops from about 16 GiB in BF16 to about 3.5 GiB. vLLM's production presets are more conservative and report 2.6x to 4.9x compression on Qwen3-4B.
Does vLLM support TurboQuant?
Yes. vLLM merged a TurboQuant attention backend in pull request 38479 on April 15, 2026. Enable it with --kv-cache-dtype turboquant_k8v4 (FP8 keys and 4-bit values, about 2.6x compression) or the more aggressive turboquant_4bit_nc, turboquant_k3v4_nc, and turboquant_3bit_nc presets. Pull request 39931, merged May 5, 2026, added hybrid-model support.
Does llama.cpp support TurboQuant?
Partly. The llama.cpp feature request for TurboQuant (issue 20977) was closed as not planned. Mainline did merge a Hadamard rotation of the KV cache before quantization (pull request 21038, April 1, 2026), which cites TurboQuant and covers its rotation stage when combined with the stock q4_0 cache. The PolarQuant codebook and the full method remain in community forks such as the TurboQuant+ project.
Is TurboQuant lossless?
TurboQuant is a lossy quantizer, but its error is small enough that Google measured no accuracy loss at about 3 bits on long-context benchmarks. On needle-in-a-haystack with Llama-3.1-8B-Instruct, TurboQuant scored 0.997, the same as the full-precision cache. Lower bit widths and real serving stacks can lose quality: vLLM's 3-bit preset scored 0.720 on GSM8K versus a 0.900 baseline on Qwen3-4B.
Related Articles
The fastest endpoints are private deployments
Morph's top speeds come from dedicated deployments, not shared public endpoints: speculators trained on your traffic, caching tuned to your workload, and volume discounts over public per-token rates. Over 500 billion tokens per day run this way.
Serve long-context models without managing the KV cache
Morph serves open models such as GLM-5.3 and DeepSeek V4.1 Flash on custom kernels behind one OpenAI-compatible API.
Sources
- Google Research: TurboQuant, Redefining AI efficiency with extreme compression (March 24, 2026)
- arXiv 2504.19874: TurboQuant, Online Vector Quantization with Near-optimal Distortion Rate
- arXiv 2502.02617: PolarQuant, Quantizing KV Caches with Polar Transformation
- arXiv 2406.03482: QJL, 1-Bit Quantized JL Transform for KV Cache Quantization with Zero Overhead
- vLLM PR 38479: TurboQuant attention backend (presets, accuracy, throughput)
- llama.cpp issue 20977: Feature Request, TurboQuant support (closed as not planned)
- Qdrant: TurboQuant in Qdrant (May 13, 2026)