Paper page - SimpleMemVLA: A Simple but Effective Native-Video Memory for Vision-Language-Action Models
2026-09-09 08:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
5mentions
SummaryAI generated
The SimpleMemVLA team has developed a visual-language-action model that does not require a dedicated memory module. This model transmits sampled history to the backbone network in timestamped video format, using only the hidden states of generated sub-tasks as the sole pathway to the standard flow-matching action heads. Since continuous decision-making shares most of the history, pre-filling shared prefixes during action execution can maintain latency at the single-frame VLA level. With fixed backbone and training settings, SimpleMemVLA outperforms retrieval, compression, and cyclic state mechanisms significantly, and causal intervention confirms that its strategy indeed reads its history. This approach has set new records in four memory benchmark tests.