Reinforcement Learning for improving Large Language Models' Catalan text simplification capabilities
2026-09-07 12:00Science🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated
A study proposes using reinforcement learning methods to improve the performance of large language models in low-resource automatic text simplification (ATS). The study introduces a new reward function that combines the SARI metric with specific penalty terms, and uses Group Relative Policy Optimization (GRPO) to guide the model in generating the desired simplified style. The experiment was validated using the IberianLLM-7B-Instruct model: the model was first trained on the English ASSET dataset, and then post-trained in two selected Catalan language benchmarks. The results showed improved performance in the ATS task, and negative behaviors observed previously were successfully suppressed. Additionally, the study attempted to perform post-training on the ASSET dataset in Catalan and Spanish through cross-language transfer learning, but no significant improvement was observed in external benchmark tests.