Predicting and Evaluating Item Responses Using Machine Learning, Text‐Embeddings, and LLMs

Abstract This study evaluates the extent to which synthetic response generation methods reproduce the psychometric properties of Likert‐type social‐emotional assessment data. We compare four approaches: machine learning (ML), text‐embedding models, zero‐shot large language models (LLMs), and fine‐tuned LLMs, in generating responses to 10 field‐test items from the Devereux Student Strengths Assessment (DESSA). Synthetic responses were evaluated using prediction accuracy, graded response model (GRM) parameter recovery, response‐style indices, and person‐level response patterns. At an aggregate level, ML and text‐embedding approaches achieved higher accuracy and more stable recovery of threshold parameters than LLM‐based methods. However, closer examination revealed systematic distortions across all methods, including shifts in scale‐use behavior and reduced variability in person‐level response patterns, particularly for LLM‐generated data. These findings suggest that while synthetic methods can approximate surface‐level response distributions, they do not consistently reproduce the response processes underlying ordinal data. Implications for the use of synthetic data in early‐stage item development and psychometric evaluation are discussed.

Authors

Institutions

Publication Details

Journal
Journal of Educational Measurement
Published
2026-10-07
DOI
https://doi.org/10.1111/jedm.70067
Primary Topic
Psychometric Methodologies and Testing
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Predicting and Evaluating Item Responses Using Machine Learning, Text‐Embeddings, and LLMs

Tong Wu, Onur Demirkaya, Hsin‐Ro Wei, Huan Liu et al.
Journal of Educational Measurement
Psychometric Methodologies and Testing
article

Predicting and Evaluating Item Responses Using Machine Learning, Text‐Embeddings, and LLMs

Tong Wu, Onur Demirkaya, Hsin‐Ro Wei, Huan Liu, Evelyn S. Johnson
article en

Abstract

Abstract This study evaluates the extent to which synthetic response generation methods reproduce the psychometric properties of Likert‐type social‐emotional assessment data. We compare four approaches: machine learning (ML), text‐embedding models, zero‐shot large language models (LLMs), and fine‐tuned LLMs, in generating responses to 10 field‐test items from the Devereux Student Strengths Assessment (DESSA). Synthetic responses were evaluated using prediction accuracy, graded response model (GRM) parameter recovery, response‐style indices, and person‐level response patterns. At an aggregate level, ML and text‐embedding approaches achieved higher accuracy and more stable recovery of threshold parameters than LLM‐based methods. However, closer examination revealed systematic distortions across all methods, including shifts in scale‐use behavior and reduced variability in person‐level response patterns, particularly for LLM‐generated data. These findings suggest that while synthetic methods can approximate surface‐level response distributions, they do not consistently reproduce the response processes underlying ordinal data. Implications for the use of synthetic data in early‐stage item development and psychometric evaluation are discussed.

Journal of Educational MeasurementVol. 63(4)
Clinical Insights (US)
Openalex Percentile: Top 9%
Psychometric Methodologies and Testing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.