AuraTracer智迹闻
中文

EVENT DOSSIER

Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model Pipelines

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

A research on grounded language-model pipelines analyzed the identity transfer situations in 600 HybridQA questions. Among 1,463 resolvable records, a hybrid retrieval strategy combining precise key-value search and title matching could retrieve the target objects 100%; however, BM25 retrieval based solely on the text led to 26.6% of records missing the objects. Although mixed retrieval improved the ranking, 1.0% of records still were missed. The study found that differences in the selected objects and the dataset’s tracking passages caused inconsistencies in 329 records, and the top five checks based on the original questions showed discrepancies in 5.9% of records. A freeze reader comparison experiment showed that the absence of aligned objects significantly reduced the precise matching score from 28.6 to 31.0 points; removing specific passages significantly decreased the precise matching rate, while removing comparable-length comparison passages did not cause such a decline.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
HybridQA

SignalsSIGNALS

Keyword heat
  • HybridQA1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model Pipelines

本研究审计了 600 个 HybridQA 问题中 grounded language-model pipelines 的身份交接情况。在 1,463 条可解析记录中,精确键值查找和精确标题匹配能 100% 找回目标对象;而仅使用正文的 BM25 检索在 389 条记录(26.6%)上遗漏该对象,混合检索加重排序仅在 14 条(1.0%)上遗漏。由于选定对象与数据集追踪 passage 身份不同导致 329 条记录出现差异,基于原始问题排名的前五名检查在 106 条(5.9%)记录中结果不一致。冻结读者对比显示,对齐对象的缺失使精确匹配分数降低 28.6 至 31.0 分;在一个精心挑选的 64 项队列中,移除该 passage 会显著降低精确匹配率,而移除长度相似的对比 passage 则无法复现此下降。研…