AI Psychometrics: Evaluating Measurement Invariance in Science Self-Efficacy Using Large Language Models

The present study (1) examined sex- and language-based measurement invariance of the science self-efficacy scale using PISA 2006 and 2015 U.S. data (N = 11,323) and (2) evaluated six large language models (ChatGPT, DeepSeek, Grok, Copilot, MetaAI, Gemini) in generating invariance approximations from prompt content and pre-existing model knowledge. Multigroup confirmatory factor analysis (MG-CFA) established empirical invariance benchmarks across sex and language groups. The same six LLMs were prompted to generate plausible fit-index values (CFI, RMSEA) and invariance decisions without access to empirical response data. Approximation discrepancy and agreement with the MG-CFA benchmarks were assessed using RMSE, MAE, Cronbach’s α and intraclass correlation coefficients (ICCs). MG-CFA supported full measurement invariance across sex and language groups in both cycles. In contrast, LLMs showed strong apparent consistency for CFI (α = 0.857) but poor convergence with empirical data (single-measure ICC = 0.302 for CFI, 0.024 for RMSEA). Discrepancies were often large enough to affect invariance interpretations and the models often overestimated model fit while producing invariance decisions that were not always aligned with the empirical MG-CFA benchmarks. Findings support the science self-efficacy scale’s cross-group validity but caution that current off-the-shelf LLMs, when prompted to generate plausible psychometric approximations without empirical data, produce unreliable outputs for high-stakes psychometric decisions. Implications, limitations and future directions are discussed.

Authors

Institutions

Publication Details

Journal
Psychology International
Published
2026-09-21
DOI
https://doi.org/10.3390/psycholint8030062
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

AI Psychometrics: Evaluating Measurement Invariance in Science Self-Efficacy Using Large Language Models

Danielle Ka Lai Lee, Onur Ramazan
Psychology International
Artificial Intelligence in Healthcare and Education
article

AI Psychometrics: Evaluating Measurement Invariance in Science Self-Efficacy Using Large Language Models

Danielle Ka Lai Lee, Onur Ramazan
article en

Abstract

The present study (1) examined sex- and language-based measurement invariance of the science self-efficacy scale using PISA 2006 and 2015 U.S. data (N = 11,323) and (2) evaluated six large language models (ChatGPT, DeepSeek, Grok, Copilot, MetaAI, Gemini) in generating invariance approximations from prompt content and pre-existing model knowledge. Multigroup confirmatory factor analysis (MG-CFA) established empirical invariance benchmarks across sex and language groups. The same six LLMs were prompted to generate plausible fit-index values (CFI, RMSEA) and invariance decisions without access to empirical response data. Approximation discrepancy and agreement with the MG-CFA benchmarks were assessed using RMSE, MAE, Cronbach’s α and intraclass correlation coefficients (ICCs). MG-CFA supported full measurement invariance across sex and language groups in both cycles. In contrast, LLMs showed strong apparent consistency for CFI (α = 0.857) but poor convergence with empirical data (single-measure ICC = 0.302 for CFI, 0.024 for RMSEA). Discrepancies were often large enough to affect invariance interpretations and the models often overestimated model fit while producing invariance decisions that were not always aligned with the empirical MG-CFA benchmarks. Findings support the science self-efficacy scale’s cross-group validity but caution that current off-the-shelf LLMs, when prompted to generate plausible psychometric approximations without empirical data, produce unreliable outputs for high-stakes psychometric decisions. Implications, limitations and future directions are discussed.

Psychology InternationalVol. 8(3)
Hong Kong Shue Yan University (CN), University of Hong Kong (HK)
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

AI Psychometrics: Evaluating Measurement Invariance in Science Self-Efficacy Using Large Language Models — Danielle Ka Lai Lee, Onur Ramazan · Psychology International (2026) | TGRS Research Map | TGRS