Quality Recovery for Quantized KV Caches via Low-Rank Attention Adaptation
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
3mentions
SummaryAI generated
To address the problem of self-regressive decoding quality loss caused by low-bit KV caches, researchers proposed updating the floating-point cache model behavior through low-rank Q/K/V projections and combining it with student-executed physical packing incremental caching. On TinyLlama-1.1B and Gemma-4-12B models, a 4-bit affine cache adapter restored a test perplexity gap of $54.24\%\pm2.47\%$ and $75.96\%\pm4.04\%$ respectively. On the frozen NF4 Llama-3.1-8B base model, the recovery rates under specific quantizers were $60.42\%$ and $37.61\%$ respectively, while maintaining 180 associated retrieval cases. Additionally, the score of the Gemma model on the official 4K/8K RULER subset increased from 42.80 before adaptation to 48.33 after adaptation.