FinalityBench: An Effect-Level Benchmark for Agent Decisions Under Delayed and Conflicting Financial Finality
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
The researchers released FinalityBench, an executable benchmark used to evaluate smart agents’ decision-making in scenarios with delayed and conflicting financial outcomes. The benchmark includes 321 tasks (including 45 twin tasks), generating disagreements through the delivery of hidden specification event logs and independent failures, and ranking them based on the economic location of merchant terminals. The tests covered over 14,445 ranked episodes from nine programmed strategies. The results showed that the “ship immediately upon first signal” strategy ranked second in single-task accuracy (65.7%), but performed worst in paired losses. The runtime-gated mechanism achieved an efficiency of 85.4% on authoritative outcome probes and did not cause additional losses like the polling strategy. Additionally, the language model could achieve the same accuracy without being informed of the strategy, but the resulting loss amount was approximately twice that of the manual gating.