Predicting and Evaluating Item Responses Using Machine Learning, Text‐Embeddings, and LLMs
Abstract This study evaluates the extent to which synthetic response generation methods reproduce the psychometric properties of Likert‐type social‐emotional assessment data. We compare four approaches: machine learning (ML), text‐embedding models, zero‐shot large language models (LLMs), and fine‐tuned LLMs, in generating responses to 10 field‐test items from the Devereux Student Strengths Assessment (DESSA). Synthetic responses were evaluated using prediction accuracy, graded response model (GRM) parameter recovery, response‐style indices, and person‐level response patterns. At an aggregate level, ML and text‐embedding approaches achieved higher accuracy and more stable recovery of threshold parameters than LLM‐based methods. However, closer examination revealed systematic distortions across all methods, including shifts in scale‐use behavior and reduced variability in person‐level response patterns, particularly for LLM‐generated data. These findings suggest that while synthetic methods can approximate surface‐level response distributions, they do not consistently reproduce the response processes underlying ordinal data. Implications for the use of synthetic data in early‐stage item development and psychometric evaluation are discussed.
Authors
- Tong Wu (ORCID: https://orcid.org/0000-0001-9708-3591)
- Onur Demirkaya (ORCID: https://orcid.org/0000-0002-3985-6485)
- Hsin‐Ro Wei
- Huan Liu
- Evelyn S. Johnson
Institutions
- Clinical Insights (US)
Publication Details
- Journal
- Journal of Educational Measurement
- Published
- 2026-10-07
- DOI
- https://doi.org/10.1111/jedm.70067
- Primary Topic
- Psychometric Methodologies and Testing
- Type
- article
- Field-Weighted Citation Impact
- 0.00