AuraTracer智迹闻
中文

EVENT DOSSIER

Reinforcement Learning for improving Large Language Models' Catalan text simplification capabilities

2026-09-07 12:00 Science 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
2mentions
SummaryAI generated

A study proposes using reinforcement learning methods to improve the performance of large language models in low-resource automatic text simplification (ATS). The study introduces a new reward function that combines the SARI metric with specific penalty terms, and uses Group Relative Policy Optimization (GRPO) to guide the model in generating the desired simplified style. The experiment was validated using the IberianLLM-7B-Instruct model: the model was first trained on the English ASSET dataset, and then post-trained in two selected Catalan language benchmarks. The results showed improved performance in the ATS task, and negative behaviors observed previously were successfully suppressed. Additionally, the study attempted to perform post-training on the ASSET dataset in Catalan and Spanish through cross-language transfer learning, but no significant improvement was observed in external benchmark tests.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
ASSETIberianLLM-7B-Instruct

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
ASSET × IberianLLM-7B-I…1

SignalsSIGNALS

Keyword heat
  • IberianLLM-7B-Instruct1
  • ASSET1

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

Reinforcement Learning for improving Large Language Models' Catalan text simplification capabilities

本文提出一种基于强化学习(RL)的方法,利用大语言模型(LLMs)提升低资源语言自动文本简化(ATS)质量。研究引入结合 SARI 指标与特定惩罚项的新型奖励函数,并通过组相对策略优化(GRPO)指导模型生成目标简化风格。该方案在 IberianLLM-7B-Instruct 模型上进行了后训练验证:基于英语 ASSET 数据集训练后,模型在两个精选的加泰罗尼亚语基准测试中 ATS 性能得到提升,并成功抑制了先前观察到的负面行为。此外,研究尝试通过跨语言迁移学习将 ASSET 翻译为加泰罗尼亚语和西班牙语进行后训练,但未能显示出对域外基准测试的显著改进效果。