Fine-Tuned LLMs Improve Imputation of Missing Survey Data Compared With Baseline
September 30, 2026
Researchers are exploring a new way to make missing survey data easier and more efficient to address, particularly when complex questionnaire skip patterns affect which questions respondents answer. A recent study, published at the 2026 Joint Statistical Meetings (JSM) proceedings, provides preliminary evidence that fine-tuning large language models (LLMs) could improve the accuracy of imputing missing responses to factual survey questions, including imputing missing responses from smaller demographic groups. Westat authors include Hanyu Sun, PhD; Daifeng Han, PhD; and J. Michael Brick, PhD.
Using 2021–2024 National Health Interview Survey (NHIS) public-use data, researchers fine-tuned Mistral-7B to identify response patterns across demographic subgroups. They tested the approach on survey questions both with and without skip logic and compared the results with those from the zero-shot baseline model that was not fine-tuned. Fine-tuning reduced differences between imputed and observed response distributions by about 90% for questions both with and without skip logic. The fine-tuned model also followed the questionnaire skip logic more consistently.
“The findings suggest that LLM fine-tuning could offer a promising approach for imputing missing survey data, including in analyses involving smaller or sparsely represented demographic groups,” notes Sun.
Learn More
Using Large Language Models for Missing Survey Data Imputation: A Proof-of-Concept
Hanyu Sun, Daifeng Han, and J. Michael Brick