AuraTracer智迹闻
中文

EVENT DOSSIER

Paper page - Group Adaptive Clipping Policy Optimization

2026-09-07 08:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
3mentions
SummaryAI generated

Sheng Jia and others proposed a new method called “Group Adaptive裁剪 Strategy Optimization (GAPO)” to improve group relative strategy optimization in reinforcement learning. This method addresses the problem of over-suppression of relevant samples caused by existing fixed importance sampling裁剪 boundaries. It solves this problem by dynamically adjusting the boundaries based on rolling advantages. GAPO does not require modifications to the reward shaping or objective function, and improved the mathematical reasoning benchmarks in pass@1 and pass@k on both Qwen and Llama-based models.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
Hugging FaceLlamaQwen

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
Hugging Face × Llama1Hugging Face × Qwen1Llama × Qwen1

SignalsSIGNALS

Keyword heat
  • Hugging Face1
  • Qwen1
  • Llama1

All reports (1)SOURCES

H Hugging Face Papers en 2026-09-07 08:00

Paper page - Group Adaptive Clipping Policy Optimization

Sheng Jia 等人提出一种名为“组自适应裁剪策略优化(GAPO)”的新方法,用于改进强化学习中的组相对策略优化。现有固定重要性采样裁剪边界导致困难问题上的正确样本被过度抑制,而 GAPO 通过根据滚动优势动态调整该边界来解决问题。该方法无需修改奖励塑造或目标函数,在 Qwen 和 Llama 基模型上均提升了数学推理基准的 pass@1 和 pass@k 表现。