AuraTracer智迹闻
中文

EVENT DOSSIER

CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated

CoSkill proposes a unified multi-agent reinforcement learning framework, aiming to achieve collaborative adaptation on a hierarchical skill library by jointly training reasoning agents and learnable meta-skills agents. This framework reconsiders static meta-skills workflows into dynamic agents, enabling the reasoning agents to take actions based on the retrieved task skills and sub-step skills, thereby guiding the meta-skills agents to optimize these skills. In the ALFWorld and WebShop benchmarks, CoSkill performed significantly better than existing baselines, achieving success rates of 98.4% and 90.6%, respectively, which is an improvement of approximately 3.5 to 6.2 percentage points compared to existing methods. Additionally, the framework demonstrated better performance in terms of sample efficiency, asymptotic performance, and wall clock efficiency. The relevant code has been made open-source.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
CoSkilljinyuan-cookie

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
CoSkill × jinyuan-cookie1

SignalsSIGNALS

Keyword heat
  • CoSkill1
  • jinyuan-cookie1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

CoSkill: Joint Reinforcement Learning of Reasoning and Meta-Skill Agents for Hierarchical Skill Evolution

CoSkill 提出一种统一的多智能体强化学习框架,将静态元技能工作流重构为可学习的元技能代理,并与推理代理在分层技能库上联合训练。该框架通过建模推理与元技能代理为共享单一骨干的合作团队,实现端到端协同适应:推理代理根据检索到的任务技能和子集步骤技能采取动作,其任务表现指导元技能代理优化这些步骤技能。在 ALFWorld 和 WebShop 实验表明,CoSkill 显著优于现有基于技能和强化学习基线,成功率达到 98.4% 和 90.6%,分别提升 3.5 和 6.2 个百分点,且在早期样本效率、渐近性能和墙钟效率方面表现更优。相关代码已开源。