ConsensusBench: Benchmark of Consensus Nodes for LLM Reasoning via Outcome Reward Densifying
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated
On September 7, 2026, ConsensusBench was released in the arXiv cs.CL domain. It aims to evaluate the reasoning consensus nodes of large language models through result-based reward densification techniques. Addressing the issue of sparse rewards in existing reinforcement learning algorithms due to their reliance on final answers, the study proposes ConsensusPR, a rule-based process-level signal approach. This method filters out correct trajectories from N samples and clusters semantically equivalent intermediate statements to identify consensus nodes, integrating them into an algorithm framework similar to GRPO. ConsensusBench introduces three evaluation metrics: accuracy of final answers, node coverage, and number of Tokens per node. Experiments show that this method outperforms existing GRPO-style methods significantly in long reasoning trajectories on datasets such as AIME 2024, AIME 2025, GSM8K, MATH-500, and ConsensusBench.