AuraTracer智迹闻
中文

EVENT DOSSIER

SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents

2026-09-07 12:00 Models 🔥 40.2 heat score
1sources
1days unfolding
40.2heat score
1mentions
SummaryAI generated

On September 7, 2026, arXiv cs.AI published the paper SiLR, which proposes a structural preservation reward method for tool agents in large language models (LLMs). This research aims to address the issue of loss of structural information during existing tool invocation processes. By designing structural preservation access and process reward mechanisms, it improves the efficiency and accuracy of LLMs in using tools for complex tasks.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
SiLR

SignalsSIGNALS

Keyword heat
  • SiLR1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

SiLR: Structure-Preserving Admission and Process Reward for LLM Tool Agents

SiLR 提出一种基于分支级违规状态乘积序的运行时门控机制,在 Gym-ANM 等基准上使多动作恢复成功率从基线 0/21 提升至 21/21。该方法通过影子执行每个提议并依据过载分支支持与单分支严重度进行判定,证明了标量代理无法准确表示该顺序,其失败源于表征而非阈值调整。在 CityLearn 及三种模型家族测试中,确定性模拟门控将 LLM 置于信任边界之外,仅全分支谓词能防御幅度重分配攻击;相比之下,标量投影和单支支持均允许物理不安全动作,其中后者占比达 63.2%。作为 GRPO 过程奖励使用时,SiLR 在所有场景下优于计数投影,且是唯一使未门控策略表现超越未训练基线(0.844 vs 0.778)的奖励函数。