AuraTracer智迹闻
中文

EVENT DOSSIER

Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models

2026-09-07 12:00 Models 🔥 40.2 heat score
1sources
1days unfolding
40.2heat score
1mentions
SummaryAI generated

On September 7, 2026, arXiv cs.LG published the paper “Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models”. This study addresses the limitations of existing diffusion models, which rely on pairwise preferences for alignment, and proposes a new paradigm of list-wise reward-aware alignment. By introducing list-formatted reward signals, this method enables the model to consider both the relative relationships between multiple samples and the overall distribution characteristics, thereby improving alignment efficiency and effectiveness while maintaining generation quality. This provides a new technical approach for optimizing diffusion models.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
Diffusion LAIR

SignalsSIGNALS

Keyword heat
  • Diffusion LAIR1

All reports (1)SOURCES

A arXiv cs.LG en 2026-09-07 12:00

Beyond Pairwise Preferences: Listwise Reward-Aware Alignment for Diffusion Models

arXiv:2605.26491v2 提出 Diffusion LAIR,一种针对扩散模型的奖励感知列表式偏好优化方法。该方法将同一提示词下候选图像的奖励分数转化为中心化优势权重,在隐式奖励(定义为当前模型相对于固定参考模型的去噪损失改进)上优化加权回归目标,并引入二次惩罚正则化其幅度。实验表明,Diffusion LAIR 在 SD1.5 和 SDXL 模型上的文本生成、组合生成及图像编辑基准测试中,均优于现有的强偏好优化基线方法。