AuraTracer智迹闻
中文

EVENT DOSSIER

Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs

2026-09-07 12:00 Models 🔥 40.2 heat score
1sources
1days unfolding
40.2heat score
1mentions
SummaryAI generated

A large-scale study on the performance of multi- and cross-lingual multiple-choice questions (MCQA) for large language models (LLMs) aims to estimate the uncertainty of the models during reasoning processes. Through extensive datasets and experiments, this study analyzes the differences in model performance and uncertainty characteristics across different linguistic contexts, providing empirical evidence for improving the reliability and interpretability of large models.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
LLMs

SignalsSIGNALS

Keyword heat
  • LLMs1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

Estimating Uncertainty from Reasoning: A Large-Scale Study of Multi- and Crosslingual MCQA Performance in LLMs

研究人员对 22 种语言中的多语种及跨语言多选择问答(MCQA)任务进行了大规模不确定性估计(UE)评估,涵盖高、中、低资源设置。研究使用两个人工筛选的问答数据集,对比了九种开放和封闭盒子 UE 方法在不同模型规模与架构下的表现,并采用长文本推理方式以消除 LLM 作为裁判及基于嵌入评分带来的噪声。主要发现包括:提示模型用英语进行推理而保留低资源语言问题能显著提升 UE 性能,表明低资源语言的阅读理解能力完好,可靠性瓶颈在于生成而非理解;用英语进行推理能有效缩小低资源与高资源语言间的 UE 性能差距,证明生成语言比问题语言更重要;UE 方法的选择应取决于模型规模,小模型中基于概率的开放盒子方法表现更优,大模型中封闭盒子自陈述不确定性方法更优越。此外,研究提供了选择性预测中的阈值选择分析,为多语种环境下的弃权校…