AuraTracer智迹闻
中文

EVENT DOSSIER

A Repeated-Measurement Study for Cultural Analytics of English Song Lyrics Using Five Large Language Models

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
0mentions
SummaryAI generated

A repeated-measurement study on cultural analysis evaluated the performance of five large language models as zero-labelgers for English lyrics, focusing on four social constructs: self-worth, self-control, a sense of belonging, and a desire for recognition. The study examined cross-run consistency and cross-model convergence by annotating the same song corpus multiple times. Results showed that self-worth measurement exhibited the strongest repeated-measure reliability among all models; the stability of the desire for recognition was poor; while self-control and the sense of belonging demonstrated moderate but model-dependent reliability. Downstream classification tasks indicated that consensus-based language model labels contained learnable signals, although transfer alone did not establish construct validity. The study suggests that results on repeated-measure stability and cross-model convergence must be reported before considering large model annotations as scalable cultural analysis measures.

Related eventsRELATED EVENTS

All reports (1)SOURCES

A arXiv cs.LG en 2026-09-07 12:00

A Repeated-Measurement Study for Cultural Analytics of English Song Lyrics Using Five Large Language Models

This study evaluated the reliability of five large language models as zero-sample annotators in measuring four social constructs in English lyrics: self-worth, self-control, a sense of belonging, and a desire for recognition. By repeatedly annotating large song corpora, the study examined three characteristics of measurements based on large language models: cross-run consistency, cross-model convergence, and the transferability of consensus labels to supervised classification. The findings showed that measurements based on large language models were not equally reliable across different constructs: self-worth exhibited the strongest repeat-measure reliability among all models, while the stability of a desire for recognition was generally poor. Self-control and a sense of belonging showed moderate but model-dependent reliability. Downstream classification further indicated that consensus large language model labels contained learnable signals, although transferability itself did not establish construct validity. Therefore, before considering large language model annotations as scalable cultural analysis measures, it is necessary to report on the repeat-measure stability and cross-model convergence.