AuraTracer智迹闻
中文

EVENT DOSSIER

Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe

2026-09-07 12:00 Science across 2 days 🔥 47.2 heat score
2sources
2days unfolding
47.2heat score
3mentions
SummaryAI generated

In September 2026, a South Korean research institution released the KOPA-Bench benchmark, which included 145 real-world tasks aimed at evaluating the performance of open-source models in multi-step tool calls across government APIs. To address the lack of existing evaluation criteria and the performance gap between models, the team proposed the EDGE method based on data synthesis from execution dynamic graphs. This method constructs a graph of tool input-output relationships and filters out successful call chains to generate executable complex trajectory data. The 9B-parameter model, fine-tuned using GRPO, performed significantly better than the unfine-tuned 27B model of the same family in both KOPA-Bench and BFCL benchmarks, achieving performance close to optimal levels.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
EDGEGRPOKorean Open Public API Benchmark

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
EDGE × GRPO2EDGE × Korean Open Publ…2GRPO × Korean Open Publ…2

Integrated timelineUNIFIED TIMELINE

  1. 2026-09-04

    Multi-Step Tool-Calling over Korean Ope…

    韩国研究机构发布多步骤工具调用基准测试 KOPA-Bench,包含 145 个真实世界任务。为解决开源模型在跨政府 API 多步调用中表现不佳且缺乏评估标准的问题,团队提出基于实时执行的动态图数据合成方法 EDGE。EDGE 构建工具输出…

  2. 2026-09-07

    Multi-Step Tool-Calling over Korean Ope…

    韩国研究人员推出 KOPA-Bench 基准测试,包含 145 个真实任务,旨在评估开源模型在跨政府 API 多步工具调用中的表现差距。为此,团队提出 EDGE(基于执行动态图的工具调用数据合成)方法,利用实时 API 调用构建并验证工具…

SignalsSIGNALS

Keyword heat
  • Korean Open Public API Benchmark2
  • EDGE2
  • GRPO2

All reports (2)SOURCES

A arXiv cs.CL en 2026-09-05 01:44

Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe

韩国研究机构发布多步骤工具调用基准测试 KOPA-Bench,包含 145 个真实世界任务。为解决开源模型在跨政府 API 多步调用中表现不佳且缺乏评估标准的问题,团队提出基于实时执行的动态图数据合成方法 EDGE。EDGE 构建工具输出与输入关联图谱,仅保留实际调用成功的链接以生成可执行的多步轨迹。经 GRPO 微调的 9B 模型在 KOPA-Bench 及 BFCL 基准测试中表现显著优于同家族未微调的 27B 模型,实现了性能逼近的效果。

A arXiv cs.AI en 2026-09-07 12:00

Multi-Step Tool-Calling over Korean Open Public APIs: A Benchmark and a Data-Synthesis Recipe

韩国研究人员推出 KOPA-Bench 基准测试,包含 145 个真实任务,旨在评估开源模型在跨政府 API 多步工具调用中的表现差距。为此,团队提出 EDGE(基于执行动态图的工具调用数据合成)方法,利用实时 API 调用构建并验证工具间的数据流转链路,进而合成可执行的复杂轨迹。该数据集经 GRPO 微调后,使 9B 参数模型性能接近同家族未调优的 27B 模型,不仅显著提升了 KOPA-Bench 表现,也在 BFCL 基准测试中取得进步。