大语言模型能否成为可靠的硅基受访者?
--基于CFPS回答分布的微调训练实验
林泽腾1叶振业2王仁和2*
(1. 香港科技大学,广东广州,511453;
2.*通讯作者,华南师范大学政治与公共管理学院,广东广州,510006)
摘要:随着生成式人工智能在社会科学研究中的应用不断扩展,大语言模型能否作为“硅基受访者”参与问卷预实验、调查设计和公众意见研究,成为计算社会科学与社会调查方法领域的重要问题。本文以中国家庭追踪调查(CFPS)2010-2022年的个人面板数据为基础,围绕生活满意度和主观健康两个主观变量,让大语言模型在学会变量分布特征的基础上,以“硅基受访者”的身份参加后续年份的问卷问答。方法上,我们采用基于Kullback-Leibler散度约束的低秩适应有监督微调(FT-KL)训练大模型,并与问答式、第一人称自传式和第三人称角色描摹式三类零样本提示词工程进行效果对比。实验结果显示,相较三类零样本提示词方法,FT-KL训练取得统计显著的改善,经过训练后的大模型可以很好地充当“硅基受访者”回答未来年份的问卷。稳健性检验进一步表明,这一结论不受主观变量设置、大模型参数规模、提示词设置的影响。研究表明,在真实的调查数据分布监督下,大模型可以成为可靠的硅基受访者。
关键词:大语言模型;硅基受访者;中国家庭追踪调查;计算社会科学
Can Large Language Models Become Reliable Silicon Respondents?
A Fine-tuning Experiment Based on CFPS Response Distributions
Lin Zeteng1Ye Zhenye2 Wang Renhe2*
(1. Hong Kong University of Science and Technology (Guangzhou), Guangzhou 511453, Guangdong, China;
2. School of Politics and Public Administration, South China Normal University, Guangzhou 510006, Guangdong, China)
Abstract: With the expanding use of generative artificial intelligence in social science research,whether large language models can serve as reliable "silicon respondents" has become an important methodological question for computational social science and survey research. Based on individual-level panel data from the China Family Panel Studies (CFPS) from 2010 to 2022, this study examineswhether large language models can predict future response distributions for two subjective variables:life satisfaction and self-rated health. Methodologically, we fine-tune large language models using supervised fine-tuning with Low-Rank Adaptation under a Kullback-Leibler divergence constraint(FT-KL), and compare this approach with three zero-shot prompting strategies: question-answer prompting, first-person biographical prompting, and third-person portrayal prompting. The resultsshow that FT-KL consistently outperforms all prompting-based baselines across different model sizes,task settings, and subjective variables. Models calibrated with real survey response distributionsgenerate predictions that are closer to actual CFPS distributions in terms of JSD, KL, L1, and EMD.These findings suggest that large language models do not naturally become reliable surveyrespondents through demographic role prompting alone. However, when constrained by real surveydistributions, they can generate statistically more realistic silicon samples within clearly defined taskboundaries. This study provides empirical evidence for using calibrated large language models as anauxiliary tool for questionnaire pretesting, policy pilot evaluation, and evidence-informed decision-making.Distribution; Computational Social Science
Keywords: Large Language Models; Silicon Respondents; China Family Panel Studies; Response
[原文下載][Download PDF]
我們將一如既往的為您提供優質的服務。
掃一掃諮詢微信客服