AuraTracer智迹闻
中文

EVENT DOSSIER

TeleTables: A Benchmark for Large Language Models in Telecom Table Interpretation

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated

The researchers released the TeleTables benchmark, which included 2,220 tables from 13 3GPP specifications in four formats, along with 500 manually verified multiple-choice questions. The evaluation showed that in closed-book mode, domain knowledge was the main limiting factor, with no general large language model achieving an accuracy of over 41%; when tables were provided as context, the best models achieved an accuracy of over 90%, but this systematically declined with increasing reasoning depth, evidence range, and structural complexity, with a 32.2 percentage point gap between different reasoning skills. The tests indicated that specialization in non-telecommunications data did not yield consistent benefits, and strong reasoning abilities were crucial for reliable interpretation of complex technical tables.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
3GPPTeleTables

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
3GPP × TeleTables1

SignalsSIGNALS

Keyword heat
  • TeleTables1
  • 3GPP1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

TeleTables: A Benchmark for Large Language Models in Telecom Table Interpretation

研究人员发布了 TeleTables,这是一个包含来自四个格式共 13 个 3GPP 规范的 2,220 张表格及 500 道经人工验证多选题的基准测试。对 20 种开源大语言模型(LLM)的评估揭示了两大性能瓶颈:在闭卷模式下,领域知识是主要限制因素,没有任何通用模型准确率超过 41%;当表格作为上下文提供时,最佳模型准确率超 90%,但随推理深度、证据范围和结构复杂度的增加而系统性下降,推理技能间存在 32.2 个百分点的差距。测试表明,在非电信数据上的专业化无一致益处,而强大的推理能力对于可靠解读复杂技术表格至关重要。