TeleTables: A Benchmark for Large Language Models in Telecom Table Interpretation
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated
The researchers released the TeleTables benchmark, which included 2,220 tables from 13 3GPP specifications in four formats, along with 500 manually verified multiple-choice questions. The evaluation showed that in closed-book mode, domain knowledge was the main limiting factor, with no general large language model achieving an accuracy of over 41%; when tables were provided as context, the best models achieved an accuracy of over 90%, but this systematically declined with increasing reasoning depth, evidence range, and structural complexity, with a 32.2 percentage point gap between different reasoning skills. The tests indicated that specialization in non-telecommunications data did not yield consistent benefits, and strong reasoning abilities were crucial for reliable interpretation of complex technical tables.