AuraTracer智迹闻
中文

EVENT DOSSIER

KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
4mentions
SummaryAI generated

The KVMem system successfully implemented virtualized working spaces with millions of tokens on consumer-grade GPUs. By storing overflow history in paginated KV states within the GPU, host memory, and NVMe, and using lightweight attention space indexing to manage query dependencies, KVMem demonstrated higher task efficiency and reasoning speed compared to mainstream compression methods in benchmark tests such as LongMemEval, MemoryAgentBench, and AgentLongBench. Particularly in the DeepSWE long context test, its task success rate increased from 43.8% with compressed management to 48.4%. Local deployment evaluations showed that the system can run the Qwen3.6/3.8-27B NVFP4 model on laptops equipped with 24GB RTX 5090 Laptop GPUs, and it supports virtualizing workloads with up to 1M tokens.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
KVMemQwen3.6/3.8-27BQwen3.8-27BRTX 5090

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
KVMem × Qwen3.6/3.8-27B1KVMem × Qwen3.8-27B1KVMem × RTX 50901Qwen3.6/3.8-27B × Qwen3…1Qwen3.6/3.8-27B × RTX 5…1Qwen3.8-27B × RTX 50901

SignalsSIGNALS

Keyword heat
  • KVMem1
  • Qwen3.8-27B1
  • Qwen3.6/3.8-27B1
  • RTX 50901

All reports (1)SOURCES

A arXiv cs.LG en 2026-09-07 12:00

KVMem: Virtualizing Million-Token Agent Workspaces on a Consumer GPU

KVMem 系统实现了在消费级 GPU 上虚拟化百万 token 代理工作空间。该系统将溢出历史以分页 KV 状态形式存储于 GPU、主机内存及 NVMe 中,利用轻量级注意力空间索引选择相关块并生成受模型原生上下文窗口限制的查询依赖执行视图。在涵盖 LongMemEval、MemoryAgentBench 和 AgentLongBench 等基准的长上下文代理测试中,KVMem 相比基于压缩的主流方法实现了更高的任务效用和推理效率;在 DeepSWE 长上下文测试中,其任务成功率从仅使用压缩管理的 43.8% 提升至 48.4%。本地部署评估显示,KVMem 可在配备 24GB RTX 5090 Laptop GPU 的笔记本电脑上运行 Qwen3.6/3.8-27B NVFP4 模型,虚拟化高达 1M t…