AuraTracer智迹闻
中文

EVENT DOSSIER

Consensus Group Relative Policy Optimization for Text Generation

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
3mentions
SummaryAI generated

The research team proposed the Consensus Group Relative Policy Optimization (C-GRPO) method, which moves the decoding technique for Minimum Bayes Risk (MBR) from the inference phase to the training phase. This method can construct the group relative objective function using only the utility function and policy samples, replacing the traditional optimization approach that relies on gold-standard references or preference data. Experiments showed that in the WMT 2024 machine translation and XSum text summarization tasks, C-GRPO achieved performance comparable to MBR decoding, while eliminating computational overhead during inference and outperforming reference-free baseline methods.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
C-GRPOGRPOMBR

Event frameEVENT FRAME

Launch

arXiv:2602.03102v2 C-GRPO 提出无需金标的新文本生成优化算法

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
C-GRPO × GRPO1C-GRPO × MBR1GRPO × MBR1

SignalsSIGNALS

Keyword heat
  • C-GRPO1
  • MBR1
  • GRPO1

All reports (1)SOURCES

A arXiv cs.LG en 2026-09-07 12:00

Consensus Group Relative Policy Optimization for Text Generation

提出 Consensus Group Relative Policy Optimization(C-GRPO)方法,将最小 Bayes 风险(MBR)解码蒸馏至训练阶段,仅需效用函数与策略样本即可实现文本生成。该方法通过构建组相对目标函数替代传统依赖金标准参考或偏好数据的推理优化模式,在 WMT 2024 机器翻译和 XSum 文本摘要实验中,实现了与 MBR 解码相当的性能且消除了推理时的计算开销,同时优于无参考基线方法。