Google's TurboQuant Compresses LLM KV Caches to Three Bits
Jen Zhu describes Google Research's TurboQuant, which compresses LLM key-value caches to 3 bits per value using random rotation and PolarQuant quantization, reporting at least 6x memory reduction and up to 8x faster attention with no measured accuracy loss. The post links to Google's research blog.
Original post · 1 min read
Google just did something even more practical for the AI era: TurboQuant compresses LLM key-value caches down to 3 bits per value using random orthogonal rotation + PolarQuant scalar quantization & optional 1-bit QJL residual correction.
=>> 6× memory reduction, up to 8× faster attention (on H100), & 0 degradation on LongBench, Needle-in-a-Haystack, and RULER for models like Gemma. No retraining, no calibration needed.
Fiction just got out-engineered by reality. 😅💚💚
Google Research @GoogleResearchIntroducing TurboQuant: Our new compression algorithm that reduces LLM key-value cache memory by at least 6x and delivers up to 8x speedup, all with zero accuracy loss, redefining AI efficiency. Read the blog to learn how it achieves these results: research.google/blog/turboquant-redefining-ai-…
