ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
4mentions
SummaryAI generated
The researchers released ER PBench, aiming to evaluate the performance of large language model (LLM) agents in business decision-making. This benchmark tested 100 fixed problems across six rounds of simulations involving pricing, production, procurement, inventory, finance, and shared market competition. The experiments were conducted in a Solo ecosystem (against opponents following fixed rules) and an Arena ecosystem (where six models competed in a shared market), resulting in 1,200 model trajectories and 7,200 decision rounds. The results showed significant differences among leading models across the competitive ecosystems: DeepSeek performed best in the Solo ecosystem, with an average estimate of 252.29M; while Gemini won in the Arena ecosystem, with an average estimate of 263.95M. Both ecosystems identified the same task-winning model in only 21 out of 100 problems.