AuraTracer智迹闻
中文

EVENT DOSSIER

ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
4mentions
SummaryAI generated

The researchers released ER PBench, aiming to evaluate the performance of large language model (LLM) agents in business decision-making. This benchmark tested 100 fixed problems across six rounds of simulations involving pricing, production, procurement, inventory, finance, and shared market competition. The experiments were conducted in a Solo ecosystem (against opponents following fixed rules) and an Arena ecosystem (where six models competed in a shared market), resulting in 1,200 model trajectories and 7,200 decision rounds. The results showed significant differences among leading models across the competitive ecosystems: DeepSeek performed best in the Solo ecosystem, with an average estimate of 252.29M; while Gemini won in the Arena ecosystem, with an average estimate of 263.95M. Both ecosystems identified the same task-winning model in only 21 out of 100 problems.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
DeepSeekERPBenchGAIR-NLPGemini

Event frameEVENT FRAME

Launch

GAIR-NLP ERPBench 发布企业决策大模型代理基准

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
DeepSeek × ERPBench1DeepSeek × GAIR-NLP1DeepSeek × Gemini1ERPBench × GAIR-NLP1ERPBench × Gemini1GAIR-NLP × Gemini1

SignalsSIGNALS

Keyword heat
  • ERPBench1
  • DeepSeek1
  • Gemini1
  • GAIR-NLP1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

ERPBench: Evaluating LLM Agents for Enterprise Decision-Making Across Competitive Market Ecologies

研究人员推出 ER PBench,这是一个用于评估企业决策代理在竞争性市场生态系统中表现的执行仪器化基准。该基准基于六轮包含定价、生产、采购、库存、财务及共享市场竞争的企业资源规划(ERP)模拟,对同一组 100 个固定问题在 Solo(对抗固定规则对手)和 Arena(六模型共享市场)两种生态中进行测试。实验涵盖六个模型家族,生成 1,200 条模型轨迹和 7,200 个决策轮次。结果显示,不同生态中的领先模型存在差异:DeepSeek 在 Solo 生态中表现最佳(平均估值 252.29M,排名 1.67),而 Gemini 在 Arena 生态中胜出(平均估值 263.95M,排名 1.76)。两种生态仅在 100 个问题中的 21 个识别出相同的任务级获胜者,Gemini 在 Arena 中的末位率从…