AuraTracer智迹闻
中文

EVENT DOSSIER

A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVR

2026-09-07 12:00 Models across 2 days 🔥 47.2 heat score
2sources
2days unfolding
47.2heat score
1mentions
SummaryAI generated

The researchers proposed an interpretable reasoning framework guided by a verifier, aimed at improving the performance of transparent educational Q&A. This framework is based on the Qwen2.5-3B-Instruct model, uses gold-anchored QLoRA for supervised adaptation, and incorporates a task-aware hybrid expert system. Through a lightweight routing mechanism, logical questions are assigned to the FOL/Z3 verifier, while physical questions are handled by a symbolic solver that understands formulas and units. In tests involving 438 examples not used in training, the introduction of Group Relative RLVR (Reasoning with Language and Verification) significantly improved the reasoning depth and interpretability metric P3, from 50.68% to 72.20%, while the combined accuracy metric P1 remained stable at 55.94%. The results show that RLVR effectively strengthened the explicit reasoning structure, and symbolic verification supplemented the reliability of neural strategies through system-level corrections, outperforming methods that relied solely on self-consistency.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
Qwen2.5-3B-Instruct

Integrated timelineUNIFIED TIMELINE

  1. 2026-09-04

    A Verifier-Guided Explainable Reasoning…

    研究人员提出了一种结合金锚定 QLoRA、任务感知混合专家及组相对 RLVR 的验证器引导可解释推理框架,用于透明教育问答。该框架利用场加权 QLoRA 监督适配 Qwen2.5-3B-Instruct,并通过轻量级路由将逻辑问题分配给 …

  2. 2026-09-07

    A Verifier-Guided Explainable Reasoning…

    研究人员提出了一种结合金锚定 QLoRA、任务感知混合专家及组相对 RLVR 的验证器引导可解释推理框架,用于透明教育问答。该框架利用场加权 QLoRA 监督将 Qwen2.5-3B-Instruct 适配至权威答案,并通过轻量级路由将逻…

SignalsSIGNALS

Keyword heat
  • Qwen2.5-3B-Instruct2

All reports (2)SOURCES

A arXiv cs.LG en 2026-09-04 22:52

A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVR

研究人员提出了一种结合金锚定 QLoRA、任务感知混合专家及组相对 RLVR 的验证器引导可解释推理框架,用于透明教育问答。该框架利用场加权 QLoRA 监督适配 Qwen2.5-3B-Instruct,并通过轻量级路由将逻辑问题分配给 FOL/Z3 验证器,物理问题分配给公式与单位感知的符号求解器。在 438 个未参与训练的示例上测试显示,RLVR 使 P3(推理深度与可解释性)从 50.68% 提升至 72.20%,而混合 P1(答案正确率)稳定在 55.94%;自一致性仅提升 P1 至 50.23%,符号验证贡献了其余混合增益。结果表明,RLVR 主要强化显式推理结构,符号验证则通过系统级保守修正补充神经策略并提高答案可靠性。

A arXiv cs.AI en 2026-09-07 12:00

A Verifier-Guided Explainable Reasoning Framework with Gold-Anchored QLoRA, Task-Aware Mixture-of-Experts, and Group-Relative RLVR

研究人员提出了一种结合金锚定 QLoRA、任务感知混合专家及组相对 RLVR 的验证器引导可解释推理框架,用于透明教育问答。该框架利用场加权 QLoRA 监督将 Qwen2.5-3B-Instruct 适配至权威答案,并通过轻量级路由将逻辑问题分配给 FOL/Z3 验证器、物理问题分配给公式与单位感知符号求解器。在 438 个保留示例测试中,RLVR 使 P3(推理深度与可解释性)从 50.68% 提升至 72.20%,混合 P1(答案正确率)稳定在 55.94%;自一致性仅提升 P1 至 50.23%,符号验证提供剩余增益。结果显示,RLVR 主要强化显式推理结构,而符号验证通过系统级修正补充神经策略以提升答案可靠性。