AuraTracer智迹闻
中文

EVENT DOSSIER

Train What You Deploy:Token-Faithful Post-Training of a Production Coding

2026-09-07 12:00 Models 🔥 42.2 heat score
1sources
1days unfolding
42.2heat score
3mentions
SummaryAI generated

To address the token and control fidelity errors present in existing code generation models during deployment, researchers proposed a fidelity-aware training coupling framework and the Certified Divergent Proximal Policy Optimization (C-DPPO) algorithm. This framework retains original prompt sampling, uses a negotiation training protocol to eliminate false model calls, and limits loss calculation to verifiable token spans; it also establishes tight bilateral total variation certification bounds, adaptive K-rules, budget-aware sequence guarantees, and fault-tolerant policy masks. Matching training and testing on the TMax-100 dataset for the Baize5B and Baize10B models showed that C-DPPO achieved consistent performance improvements of 3.0 points across model sizes compared to standard DPPO. Certificate audits verified the reliability and full-featured coverage of this certification training pipeline.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
Baize10BBaize5BTMax-100

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
Baize10B × Baize5B1Baize10B × TMax-1001Baize5B × TMax-1001

SignalsSIGNALS

Keyword heat
  • Baize5B1
  • Baize10B1
  • TMax-1001

All reports (1)SOURCES

A arXiv cs.AI en 2026-09-07 12:00

Train What You Deploy:Token-Faithful Post-Training of a Production Coding

研究人员提出一种保真度感知训练耦合框架及认证发散近端策略优化(C-DPPO)算法,旨在解决现有代码与终端代理后训练流程中存在的严重令牌与控制保真度错误问题。该框架通过保留原始提示词采样、利用协商训练协议消除虚假模型调用,并将损失计算限制在可验证的令牌跨度内;C-DPPO 在此基础上建立了紧密的双边总变差认证界限、自适应 K 规则、预算感知序列保证及容错策略掩码。在 TMax-100 数据集上对 Baize5B 和 Baize10B 模型进行匹配训练与测试时,C-DPPO 相比标准 DPPO 在各模型规模下均获得一致的 3.0 分性能提升,证书审计验证了该认证训练管道的可靠性与全功能覆盖。