You Really Didn't Get That? Benchmarking Social Pragmatic Inference for Indirect and Playful Chinese Online Comments
Based on over 200,000 publicly available records of Chinese social media interactions, researchers constructed a benchmark containing 4,735 manually verified diagnostic items to evaluate the ability of large language models (LLMs) to recover indirect and jocular social meanings in specific conversations. This benchmark paired target comments with reconstructed pre-contexts and possible misinterpretations, and tested the performance of eight LLMs as questioners and answerers in a cross-writing setting. The results showed that the best model’s accuracy outside the context of the original posts was only 81.42%, while the average accuracy of all models was 68.70%, compared to 90.8% for humans. Case analysis indicated that models could often identify broad forms of irony or humor, but often misjudged its specific mechanisms or interactive behaviors.