LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
On September 7, 2026, the paper “LookStep” was published on arXiv cs.CV. This study proposes a unified end-to-end framework that combines language center future state modeling with event-driven rolling memory, aiming to improve resource efficiency in visual-linguistic navigation. The method uses language tags to generate coarse-grained navigation progress and future states for candidate actions, and autonomously decides when to write observations into bounded rolling memory. In the VLN-CE task, LookStep outperforms existing methods under the same training settings, achieving a success rate of 49.7% on the R2R-CE Val-Unseen dataset, while also offering better memory efficiency and lower data usage.