Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matching
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
On September 7, 2026, arXiv cs.CV published the paper “Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matching”, which proposes a unified single-stage optimization framework based on sample-guided distribution matching (DM-Align). This research aims to address the issues of high computational costs and model collapse in video generation using reinforcement learning. The method utilizes the concepts of DPO and GRPO to explore and construct distribution differences from preference pairs or within groups, directly constructing gradient directions that point towards human-preferred samples. By combining the gradients of real and fake models and preference-guided gradients, it eliminates the need for multi-step reward evaluation and complex ODE-SDE transformations in traditional reinforcement learning. Comprehensive experiments across multiple basic video models show that this framework outperforms independent variants and two-stage pipeline schemes in terms of distillation quality and preference alignment.
A unified single-stage optimization framework based on sample-guided distribution matching is proposed to address the issues of high computational costs and model collapse in video generation using reinforcement learning. This method introduces DM-Align, utilizing DPO and GRPO concepts to explore and construct distribution differences within preference pairs or groups, and directly generate gradient directions towards human-preferred samples. By coordinating the gradients of real and fake models as well as preference-guided gradients, the requirements for multi-step reward evaluation and complex ODE-SDE transformations in traditional reinforcement learning are eliminated. Comprehensive experiments across multiple base video models show that this framework outperforms independent variants and two-stage pipeline solutions in terms of distillation quality and preference alignment.