Generating Constructive Feedback on Stories via Reinforcement Learning
研究人员提出一种强化学习方法,利用组相对策略优化(GRPO)引导大语言模型生成建设性反馈,无需真实反馈数据。该方法采用新颖的多组件奖励函数,优先选择针对故事独特、能提升质量并解决关键写作问题的反馈。在三个故事语料库的自动与人工评估中,该方案表现优于包括 Gemini 在内的最先进大语言模型及竞争性基线。研究证实,提供可操作建议是驱动反馈建设性的主要因素。
EVENT DOSSIER
The researchers proposed a method based on reinforcement learning (RL), using Group Relative Policy Optimization (GRPO) to guide large language models in providing constructive feedback for stories. This approach innovatively uses multi-component reward functions, prioritizing feedback that improves quality, addresses key writing issues, and is tailored to the uniqueness of the story, without relying on real-world feedback data. In automatic and manual evaluations conducted on three story corpora, this approach performed better than state-of-the-art large language models, including Gemini, and competitive baselines. The study confirmed that providing actionable advice is the main factor driving constructive feedback.
研究人员提出一种强化学习方法,利用组相对策略优化(GRPO)引导大语言模型生成建设性反馈,无需真实反馈数据。该方法采用新颖的多组件奖励函数,优先选择针对故事独特、能提升质量并解决关键写作问题的反馈。在三个故事语料库的自动与人工评估中,该方案表现优于包括 Gemini 在内的最先进大语言模型及竞争性基线。研究证实,提供可操作建议是驱动反馈建设性的主要因素。