AuraTracer智迹闻
中文

EVENT DOSSIER

LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

On September 7, 2026, the paper “LookStep” was published on arXiv cs.CV. This study proposes a unified end-to-end framework that combines language center future state modeling with event-driven rolling memory, aiming to improve resource efficiency in visual-linguistic navigation. The method uses language tags to generate coarse-grained navigation progress and future states for candidate actions, and autonomously decides when to write observations into bounded rolling memory. In the VLN-CE task, LookStep outperforms existing methods under the same training settings, achieving a success rate of 49.7% on the R2R-CE Val-Unseen dataset, while also offering better memory efficiency and lower data usage.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
LookStep

SignalsSIGNALS

Keyword heat
  • LookStep1

All reports (1)SOURCES

A arXiv cs.CV en 2026-09-07 12:00

LookStep: Efficient Vision-Language Navigation with Linguistic Foresight and Event Driven Memory

LookStep 提出一种结合语言中心未来状态建模与事件驱动滚动记忆的统一端到端框架,旨在提升视觉 - 语言导航的资源效率。该方法利用语言标签为候选动作生成粗粒度导航进度和未来状态,并自主决定将观察写入有界滚动记忆。在 VLN-CE 任务中,LookStep 在同训练设置下优于现有方法,在 R2R-CE Val-Unseen 数据集上取得 49.7% 的成功率,同时具备更好的内存效率和更低的数据使用量。