dorsal/arxiv
View SchemaKVzap: Fast, Adaptive, and Faithful KV Cache Pruning
| Authors | Simon Jegou, Maximilian Jeblick |
|---|---|
| Categories | |
| ArXiv ID | 2601.07891vv1 |
| URL | https://arxiv.org/abs/2601.07891 |
| License | http://creativecommons.org/licenses/by/4.0/ |
Abstract
Growing context lengths in transformer-based language models have made the key-value (KV) cache a critical inference bottleneck. While many KV cache pruning methods have been proposed, they have not yet been adopted in major inference engines due to speed--accuracy trade-offs. We introduce KVzap, a fast, input-adaptive approximation of KVzip that works in both prefilling and decoding. On Qwen3-8B, Llama-3.1-8B-Instruct, and Qwen3-32B across long-context and reasoning tasks, KVzap achieves $2$--$4\times$ KV cache compression with negligible accuracy loss and achieves state-of-the-art performance on the KVpress leaderboard. Code and models are available at https://github.com/NVIDIA/kvpress.
{
"annotation_id": "fb77d63b-2ee9-4e7d-a4f4-4a12e8bfd0cc",
"date_created": "2026-02-17T05:53:12.331000Z",
"date_modified": "2026-02-17T05:53:12.331000Z",
"file_hash": "7fd35e6c55e058cb799ca675747a59bf48c471e10658c550dd8a131f2d7871f1",
"private": false,
"record": {
"abstract": "Growing context lengths in transformer-based language models have made the key-value (KV) cache a critical inference bottleneck. While many KV cache pruning methods have been proposed, they have not yet been adopted in major inference engines due to speed--accuracy trade-offs. We introduce KVzap, a fast, input-adaptive approximation of KVzip that works in both prefilling and decoding. On Qwen3-8B, Llama-3.1-8B-Instruct, and Qwen3-32B across long-context and reasoning tasks, KVzap achieves $2$--$4\\times$ KV cache compression with negligible accuracy loss and achieves state-of-the-art performance on the KVpress leaderboard. Code and models are available at https://github.com/NVIDIA/kvpress.",
"arxiv_id": "2601.07891",
"authors": [
"Simon Jegou",
"Maximilian Jeblick"
],
"categories": [
"cs.LG",
"cs.AI",
"cs.CL"
],
"license": "http://creativecommons.org/licenses/by/4.0/",
"title": "KVzap: Fast, Adaptive, and Faithful KV Cache Pruning",
"url": "https://arxiv.org/abs/2601.07891",
"version": "v1"
},
"schema_id": "dorsal/arxiv",
"source": {
"execution_id": "6debbc8c-2e9a-4bdd-b998-afe453050f06",
"id": "arXiv Dataset",
"type": "Model",
"variant": "snapshot-2026-01-17",
"version": "0.1.0"
},
"user_id": 1000002
}