AuraTracer智迹闻
中文

EVENT DOSSIER

GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated

On September 7, 2026, arXiv published the GPTNT benchmark report, evaluating the real-time collaboration capabilities of multimodal agents in the game “Keep Talking And Nobody Explodes”. The test required two agents to communicate asynchronously without turn-taking restrictions, information asymmetry, or time pressure to remove bombs. Results showed that none of the open-source and closed-source models succeeded in completing the task within the countdown. The study pointed out that existing models had critical weaknesses in state tracking, action efficiency, ambiguity handling, and error recovery. GPTNT utilized the features of game-programmed generation and an active mod community to ensure that the benchmarks continued to evolve as model capabilities improved.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
GPTNTKeep Talking And Nobody Explodes

Event frameEVENT FRAME

Launch

arXiv:2606.28514v2 GPTNT 基于游戏构建的多模态代理实时协作基准

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
GPTNT × Keep Talking An…1

SignalsSIGNALS

Keyword heat
  • GPTNT1
  • Keep Talking And Nobody Explodes1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

GPTNT: Benchmarking Real-Time Collaboration Between Multimodal Agents on Keep Talking And Nobody Explodes

GPTNT 基准测试在《Keep Talking And Nobody Explodes》游戏中评估多模体智能体实时协作能力,结果显示所有测试的开源及闭源模型均未在倒计时内成功拆除炸弹。该基准基于游戏机制构建,要求两名智能体分别持有炸弹和拆弹指令,必须在无轮替限制、信息不对称和时间压力的条件下进行异步实时沟通。研究通过控制实验发现,现有模型在状态跟踪、时间预算内的行动效率、歧义处理及错误恢复方面存在关键弱点。GPTNT 利用游戏程序化生成特性及活跃模组社区,确保基准随模型能力提升而持续演化,避免被一次性破解。