AuraTracer智迹闻
中文

EVENT DOSSIER

ConsensusBench: Benchmark of Consensus Nodes for LLM Reasoning via Outcome Reward Densifying

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated

On September 7, 2026, ConsensusBench was released in the arXiv cs.CL domain. It aims to evaluate the reasoning consensus nodes of large language models through result-based reward densification techniques. Addressing the issue of sparse rewards in existing reinforcement learning algorithms due to their reliance on final answers, the study proposes ConsensusPR, a rule-based process-level signal approach. This method filters out correct trajectories from N samples and clusters semantically equivalent intermediate statements to identify consensus nodes, integrating them into an algorithm framework similar to GRPO. ConsensusBench introduces three evaluation metrics: accuracy of final answers, node coverage, and number of Tokens per node. Experiments show that this method outperforms existing GRPO-style methods significantly in long reasoning trajectories on datasets such as AIME 2024, AIME 2025, GSM8K, MATH-500, and ConsensusBench.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
ConsensusBenchGRPO

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
ConsensusBench × GRPO1

SignalsSIGNALS

Keyword heat
  • ConsensusBench1
  • GRPO1

All reports (1)SOURCES

A arXiv cs.CL en 2026-09-07 12:00

ConsensusBench: Benchmark of Consensus Nodes for LLM Reasoning via Outcome Reward Densifying

ConsensusBench 发布,旨在通过结果奖励稠密化评估大语言模型推理共识节点。该研究针对现有强化学习算法仅依赖最终答案导致奖励稀疏的问题,提出基于规则的过程级信号 ConsensusPR。该方法从 N 次采样中筛选正确轨迹并聚类语义等价中间陈述以识别共识节点,将其整合进 GRPO 风格算法中。ConsensusBench 引入最终答案准确率、节点覆盖率和每节点 Token 数三项指标,并在 AIME 2024、AIME 2025、GSM8K、MATH-500 及 ConsensusBench 数据集上的实验表明,该方法在长推理轨迹中显著优于现有 GRPO 风格方法。