AuraTracer智迹闻
中文

EVENT DOSSIER

Unifying ICL, SFT, KL-Regularized RL Through a Bayesian Lens

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
1mentions
SummaryAI generated

On September 7, 2026, arXiv published a paper proposing a unified framework based on Bayesian perspectives, aimed at integrating supervised fine-tuning (SFT), few-shot context learning (ICL), and KL-regularized reinforcement learning (RL). The study formalized these three paradigms as摊销 and projection operations for Bayesian posterior predictions. The method consists of two steps: first, constructing a generalized Bayesian posterior using prior models and utility signals; second, approximating it as a parameterized distribution through forward KL projection, corresponding to weights within the context (SFT/RL) or within the context itself (ICL). The paper further demonstrated that KL-regularized RLHF/RLVR, reward-weighted SFT/ICL, and advantage-weighted SFT all involve forward KL projection for posterior predictions induced by rewards or advantages. Additionally, the article explored the implications of this framework for modern inference pipelines, including treating RLHF/RLVR as a “posterior design plus projection” process, and emphasizing the necessity of cold start or supervised预热 for weighted KL projection, while also mentioning…

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
DeepSeek-R1

SignalsSIGNALS

Keyword heat
  • DeepSeek-R11

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

Unifying ICL, SFT, KL-Regularized RL Through a Bayesian Lens

本文提出一种贝叶斯视角,将监督微调(SFT)、少样本上下文学习(ICL)及 KL 正则化强化学习等范式统一。核心方法包含两步:(i) 利用先验模型与效用信号构建广义贝叶斯后验;(ii) 通过前向 KL 投影将其近似为参数化分布,分别对应权重内(SFT/RL)或上下文内(ICL)。文章第一部分形式化了少样本 ICL 和 SFT 为对贝叶斯后验预测的摊销与权重内投影;第二部分至第四部分证明 KL 正则化 RLHF/RLVR、奖励加权 SFT/ICL 及优势加权 SFT 均为针对奖励或优势诱导后验的前向 KL 投影。第五部分探讨了其对现代推理管道的启示,包括将 RLHF/RLVR 视为“后验设计 + 投影”、强调冷启动或监督预热对重要性加权 KL 投影的必要性,以及 DeepSeek-R1 和 o1 风格模型结合测…