AuraTracer智迹闻
中文

EVENT DOSSIER

TokenDial: Continuous Attribute Control for Text-to-Video Generation in Visual Dial Space

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

On September 7, 2026, arXiv released the TokenDial framework, aimed at solving the problem of continuous attribute control in text-generated videos. This method constructs a Visual Dial Space V+ at the dimension of visual patch channels in the visual diffusion Transformer, and controls it through broadcast addition. TokenDial freezes the pre-trained video generator and optimizes only the addition direction, using the effect of editing videos along target attributes while keeping the rest stable, without the need for paired editing videos. The learned direction becomes a reusable visual knob, supporting continuous appearance and motion control, explicit spatiotemporal positioning, and reuse across prompts, resolutions, and video lengths. Experiments and manual evaluations show that TokenDial outperforms existing video editing and slider methods in terms of slider controllability and content retention.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
TokenDial

Event frameEVENT FRAME

Launch

arXiv:2603.27520v2 TokenDial 提出视觉拨号空间连续属性控制框架

SignalsSIGNALS

Keyword heat
  • TokenDial1

All reports (1)SOURCES

A arXiv cs.CV en 2026-09-07 12:00

TokenDial: Continuous Attribute Control for Text-to-Video Generation in Visual Dial Space

arXiv:2603.27520v2 发布 TokenDial,提出在视频扩散 Transformer 的视觉补丁通道维度构建 Visual Dial Space V+,通过广播加法方向控制生成视频的外观或运动属性。该框架冻结预训练视频生成器,仅优化加法方向,利用编辑视频沿目标属性移动且其余内容稳定的效果进行监督,无需配对编辑视频。学习到的方向成为可复用的视觉旋钮,支持连续外观与运动控制、显式时空定位及跨提示词、分辨率和视频长度的复用。实验与人工评估显示,TokenDial 在滑块可控性和内容保持方面优于现有视频编辑和滑块方法。