TokenDial: Continuous Attribute Control for Text-to-Video Generation in Visual Dial Space
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
On September 7, 2026, arXiv released the TokenDial framework, aimed at solving the problem of continuous attribute control in text-generated videos. This method constructs a Visual Dial Space V+ at the dimension of visual patch channels in the visual diffusion Transformer, and controls it through broadcast addition. TokenDial freezes the pre-trained video generator and optimizes only the addition direction, using the effect of editing videos along target attributes while keeping the rest stable, without the need for paired editing videos. The learned direction becomes a reusable visual knob, supporting continuous appearance and motion control, explicit spatiotemporal positioning, and reuse across prompts, resolutions, and video lengths. Experiments and manual evaluations show that TokenDial outperforms existing video editing and slider methods in terms of slider controllability and content retention.