AuraTracer智迹闻
中文

EVENT DOSSIER

TACIT-Switch: Cost-Aware Model Escalation for LLM Agents from Censored Supervision

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

The researchers proposed the TACIT-Switch method, which uses Teacher-Annotated Censored Intervention Times (TACIT) data marked as interval-truncated observations on a cumulative risk scale. This method learns permanent transition strategies through a hybrid-cure threshold model, achieving cost-perception improvement of language model agents without teacher involvement in deployment. Experiments show that this method treats annotations as interval-truncated observations on a cumulative risk scale and estimates the probability of strong rolling success and conditional transition thresholds. In mechanism-based multi-step simulations, TACIT-Switch achieves 7.4–11.1 percentage points higher success rate and comparable cost compared to task-level, step-level, and fixed-prefix routing baselines. Ablation experiments reveal that task characteristics and cumulative trajectory risk provide complementary information. After selecting working points based on development data, this method achieves 48.5% for the 4B Cheap model and 45.5% for the 9B Cheap model in ALFWorld.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
TACIT-SWITCH

SignalsSIGNALS

Keyword heat
  • TACIT-SWITCH1

All reports (1)SOURCES

A arXiv cs.LG en 2026-09-07 12:00

TACIT-Switch: Cost-Aware Model Escalation for LLM Agents from Censored Supervision

TACIT-SWITCH 方法利用 Teacher-Annotated Censored Intervention Times(TACIT)数据,通过混合 - 治愈阈值模型学习永久交接策略,在无需教师参与部署的情况下实现语言模型代理的成本感知升级。该方法将标注表示为累积风险尺度上的区间截断观测,并估计强滚动成功的概率及条件交接阈值。在机制式多步模拟中,TACIT-SWITCH 相比任务级、步骤级及固定前缀路由基线,成功率高出 7.4-11.1 个百分点且成本相当;消融实验表明任务特征与累积轨迹风险提供互补信息。基于开发数据选择工作点后,该方法在 ALFWorld(4B Cheap 模型达 48.5%,9B Cheap 模型达 45.5%)和 DABench(73.1%)上均实现了学习策略中最高的留样成功率。