AuraTracer智迹闻
中文

EVENT DOSSIER

Quality Recovery for Quantized KV Caches via Low-Rank Attention Adaptation

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
3mentions
SummaryAI generated

To address the problem of self-regressive decoding quality loss caused by low-bit KV caches, researchers proposed updating the floating-point cache model behavior through low-rank Q/K/V projections and combining it with student-executed physical packing incremental caching. On TinyLlama-1.1B and Gemma-4-12B models, a 4-bit affine cache adapter restored a test perplexity gap of $54.24\%\pm2.47\%$ and $75.96\%\pm4.04\%$ respectively. On the frozen NF4 Llama-3.1-8B base model, the recovery rates under specific quantizers were $60.42\%$ and $37.61\%$ respectively, while maintaining 180 associated retrieval cases. Additionally, the score of the Gemma model on the official 4K/8K RULER subset increased from 42.80 before adaptation to 48.33 after adaptation.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
Gemma-4-12BLlama-3.1-8BTinyLlama-1.1B

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
Gemma-4-12B × Llama-3.1…1Gemma-4-12B × TinyLlama…1Llama-3.1-8B × TinyLlam…1

SignalsSIGNALS

Keyword heat
  • TinyLlama-1.1B1
  • Gemma-4-12B1
  • Llama-3.1-8B1

All reports (1)SOURCES

A arXiv cs.CL en 2026-09-07 12:00

Quality Recovery for Quantized KV Caches via Low-Rank Attention Adaptation

低比特键值缓存(KV Caches)虽能减少自回归解码内存占用,但会导致质量损失。本研究提出通过低秩 Q/K/V 投影更新蒸馏浮点缓存模型行为,同时学生执行物理打包增量缓存。在 TinyLlama-1.1B 和 Gemma-4-12B 上,4 位仿射缓存适配器分别恢复 $54.24\%\pm2.47\%$ 和 $75.96\%\pm4.04\%$ 的测试 perplexity 差距;在冻结的 NF4 Llama-3.1-8B 基座上,特定量化器下恢复率分别为 $60.42\%$ 和 $37.61\%$,同时保持 180 个关联检索案例。Gemma 在官方 4K/8K RULER 子集上的分数从未适配的 42.80 提升至适配后的 48.33。此外,2 位秩 - 令牌扫描将 TinyLlama 的 2 位 pe…