A Repeated-Measurement Study for Cultural Analytics of English Song Lyrics Using Five Large Language Models
This study evaluated the reliability of five large language models as zero-sample annotators in measuring four social constructs in English lyrics: self-worth, self-control, a sense of belonging, and a desire for recognition. By repeatedly annotating large song corpora, the study examined three characteristics of measurements based on large language models: cross-run consistency, cross-model convergence, and the transferability of consensus labels to supervised classification. The findings showed that measurements based on large language models were not equally reliable across different constructs: self-worth exhibited the strongest repeat-measure reliability among all models, while the stability of a desire for recognition was generally poor. Self-control and a sense of belonging showed moderate but model-dependent reliability. Downstream classification further indicated that consensus large language model labels contained learnable signals, although transferability itself did not establish construct validity. Therefore, before considering large language model annotations as scalable cultural analysis measures, it is necessary to report on the repeat-measure stability and cross-model convergence.