AuraTracer智迹闻
中文

EVENT DOSSIER

Can Large Language Models Anticipate Behavioral Responses to Social Policies? A Case of Pension Enrollment Prediction among China's Flexible Workers

2026-09-07 12:00 Models across 2 days 🔥 47.2 heat score
2sources
2days unfolding
47.2heat score
5mentions
SummaryAI generated

On September 7, 2026, researchers published the first dedicated model for predicting pension participation among flexible workers in China, FlexPension-LLM, on arXiv. This model uses the DKI-RDistill method to inject policy-based prompts containing Probit marginal effects and household registration province rules. It utilizes LoRA/SFT to enhance supervised distillation into an open-source mixed expert (MoE) student model, with teacher errors corrected by real labels. In the CHFS 2019 blind test, FlexPension-LLM achieved a composite F1 score of 0.9316, surpassing its Claude Sonnet 4.5 teacher and 17 baseline models, and showing no significant statistical difference from Claude Opus 4.6; in four external surveys, it averaged a composite F1 of 0.7549, with the narrowest performance range. Analysis indicates that the performance improvement is mainly due to policy-based prompt injection and error correction mechanisms.

Related eventsRELATED EVENTS
Key entitiesKEY ENTITIES
CHFSClaude Opus 4.6Claude Sonnet 4.5DKI-RDistillFlexPension-LLM

Coverage · reports per dayLANGUAGE SPLIT

Entity relations
CHFS × DKI-RDistill2CHFS × FlexPension-LLM2DKI-RDistill × FlexPens…2CHFS × Claude Opus 4.61CHFS × Claude Sonnet 4.51Claude Opus 4.6 × Claud…1

Integrated timelineUNIFIED TIMELINE

  1. 2026-09-04

    Can Large Language Models Anticipate Be…

    研究人员提出使用大型语言模型(LLMs)作为政策评估工具,并发布了首个针对中国灵活就业人员层级养老金参保预测的领域专用模型 FlexPension-LLM。该模型通过 DKI-RDistill 注入包含 Probit 边际效应和户籍省份养…

  2. 2026-09-07

    Can Large Language Models Anticipate Be…

    本文提出将大型语言模型(LLMs)作为政策评估工具,并发布了首个针对中国灵活就业人员层级养老金参保预测的领域专用模型 FlexPension-LLM。该模型通过 DKI-RDistill 方法注入包含 Probit 边际效应和户籍省份养老…

SignalsSIGNALS

Keyword heat
  • FlexPension-LLM2
  • DKI-RDistill2
  • CHFS2
  • Claude Sonnet 4.51
  • Claude Opus 4.61

All reports (2)SOURCES

A arXiv cs.CL en 2026-09-04 22:24

Can Large Language Models Anticipate Behavioral Responses to Social Policies? A Case of Pension Enrollment Prediction among China's Flexible Workers

研究人员提出使用大型语言模型(LLMs)作为政策评估工具,并发布了首个针对中国灵活就业人员层级养老金参保预测的领域专用模型 FlexPension-LLM。该模型通过 DKI-RDistill 注入包含 Probit 边际效应和户籍省份养老金规则的提示词,利用 LoRA/SFT 将增强监督蒸馏至开源混合专家(MoE)学生模型中,并由真实标签修正教师错误。在 CHFS 2019 盲测中,FlexPension-LLM 实现 0.9316 的复合 F1 分数,超越其 Claude Sonnet 4.5 教师及 17 个基线中的 15 个,且与 Claude Opus 4.6 统计上无显著差异;在四次外部调查中平均复合 F1 为 0.7549。分析表明,性能提升主要源于政策提示词注入和错误过滤监督,而推理过程提供了…

A arXiv cs.CL en 2026-09-07 12:00

Can Large Language Models Anticipate Behavioral Responses to Social Policies? A Case of Pension Enrollment Prediction among China's Flexible Workers

本文提出将大型语言模型(LLMs)作为政策评估工具,并发布了首个针对中国灵活就业人员层级养老金参保预测的领域专用模型 FlexPension-LLM。该模型通过 DKI-RDistill 方法注入包含 Probit 边际效应和户籍省份养老金规则的基于政策的提示词,利用 LoRA/SFT 将增强监督蒸馏至开源混合专家(MoE)学生模型中,并由真实标签修正教师错误。在 CHFS 2019 盲测中,FlexPension-LLM 实现 0.9316 的复合 F1 分数,超越其 Claude Sonnet 4.5 教师及 17 个基线中的 15 个,且与 Claude Opus 4.6 统计无显著差异;在四项外部调查中平均复合 F1 为 0.7549,表现范围最窄。分析表明,性能提升主要源于基于政策的提示词注入和错误…