A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models
This study conducted a unified evaluation of methods for enhancing cross-lingual consistency in multilingual models, covering reasoning interventions and post-training approaches. The results showed that direct distribution alignment improved cross-lingual consistency across all model-dataset combinations, while other methods were significantly affected by answer format and language coverage breadth; cross-domain transfer effects were limited unless the output formats of the source and target tasks were similar. In cultural diversity question-answering benchmarks, no systematic decline in controlled closed-form evaluations was observed, but open-ended generation occasionally showed a decrease in accuracy for non-English answers.