Does the Selected Object Reach the Reader? Auditing Identity Handoffs in Grounded Language-Model Pipelines
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated
A research on grounded language-model pipelines analyzed the identity transfer situations in 600 HybridQA questions. Among 1,463 resolvable records, a hybrid retrieval strategy combining precise key-value search and title matching could retrieve the target objects 100%; however, BM25 retrieval based solely on the text led to 26.6% of records missing the objects. Although mixed retrieval improved the ranking, 1.0% of records still were missed. The study found that differences in the selected objects and the dataset’s tracking passages caused inconsistencies in 329 records, and the top five checks based on the original questions showed discrepancies in 5.9% of records. A freeze reader comparison experiment showed that the absence of aligned objects significantly reduced the precise matching score from 28.6 to 31.0 points; removing specific passages significantly decreased the precise matching rate, while removing comparable-length comparison passages did not cause such a decline.