Reward-Aware Trajectory Shaping for Few-step Visual Generation
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
On September 7, 2026, an arXiv cs.CV preprint arXiv:2604.14910v4 proposed a new framework called Reward-Aware Trajectory Shaping (RATS). This study aims to address the limitations of step-less visual generation models, which are constrained by multi-step teacher models. RATS optimizes performance by aligning the latent trajectories of teachers and students during the key denoising stage and introducing a reward-aware gating mechanism based on relative reward performance. Specifically, when the teacher model performs better, the method strengthens the shaping of the trajectory; when the student model’s performance matches or exceeds that of the teacher, the constraints are relaxed for continuous optimization. Experimental results show that RATS achieves preference knowledge transfer without additional computational overhead for testing time, significantly improving the balance between efficiency and quality in step-less generation and greatly narrowing the performance gap between step-less student models and strong multi-step generators.