Consensus Group Relative Policy Optimization for Text Generation
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
3mentions
SummaryAI generated
The research team proposed the Consensus Group Relative Policy Optimization (C-GRPO) method, which moves the decoding technique for Minimum Bayes Risk (MBR) from the inference phase to the training phase. This method can construct the group relative objective function using only the utility function and policy samples, replacing the traditional optimization approach that relies on gold-standard references or preference data. Experiments showed that in the WMT 2024 machine translation and XSum text summarization tasks, C-GRPO achieved performance comparable to MBR decoding, while eliminating computational overhead during inference and outperforming reference-free baseline methods.