AuraTracer智迹闻
中文

EVENT DOSSIER

FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

The researchers released FinalityBench, an executable benchmark used to evaluate smart agents’ decision-making in scenarios with delayed and conflicting financial outcomes. The benchmark includes 321 tasks (including 45 twin tasks), generating disagreements through the delivery of hidden specification event logs and independent failures, and ranking them based on the economic location of merchant terminals. The tests covered over 14,445 ranked episodes from nine programmed strategies. The results showed that the “ship immediately upon first signal” strategy ranked second in single-task accuracy (65.7%), but performed worst in paired losses. The runtime-gated mechanism achieved an efficiency of 85.4% on authoritative outcome probes and did not cause additional losses like the polling strategy. Additionally, the language model could achieve the same accuracy without being informed of the strategy, but the resulting loss amount was approximately twice that of the manual gating.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
FinalityBench

SignalsSIGNALS

Keyword heat
  • FinalityBench1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality

研究人员发布 FinalityBench,这是一个用于评估代理在延迟和冲突金融最终性下决策的可执行基准。该基准包含 321 个任务(含 45 对孪生任务),通过隐藏规范事件日志与独立故障的交付流生成分歧,依据商户终端经济位置进行分级。测试涵盖来自九种程序化策略的超过 14,445 个分级剧集,其中“首次信号即发货”策略在单任务准确率上排名第二(65.7%),但在配对损失中表现最差。运行时门控机制在权威最终性探针上达到 85.4% 效率,且未像轮询策略那样产生额外损失;语言模型在未被告知该策略的情况下,能以相同准确率实现效果,但损失金额约为人工门控的两倍。