Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning
2026-09-07 12:00Models🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
3mentions
SummaryAI generated
The researchers evaluated three open-source large language models from the United States, Poland, and China: Gemma3-12B, Bielik-11B-v3, and Qwen3-4B. They found that Qwen3-4B performed the worst on the Chinese population in its home country, with the highest attribution degree. Through target LoRA fine-tuning for five worst-case populations (requiring less than 1,200 pairs of training data and taking less than 15 minutes on a single GPU), Bielik-11B’s bias was reduced by 16.8% (p_Bonf = 0.002). However, breakdown by country showed that fine-tuning only redistributed the bias rather than eliminating it: Bielik’s worst-case population completely shifted from the American elderly to the Chinese elderly, with no overlap between the pre- and post-fine-tuning correction sets. This is the first study known to perform LoRA fine-tuning on worst-case demographic characteristics to mitigate cross-cultural biases.