Refining the Immersive Technology Evaluation Measure for Virtual Reality Clinical Skills Among Native Arabic-Speaking Medical Students: Qualitative Study

Abstract Background Extended reality technologies, including virtual reality (VR), augmented reality, and mixed reality, are increasingly used in medical education to create immersive and interactive learning environments. As these modalities expand, validated instruments are needed to measure learners’ experiences accurately. The Immersive Technology Evaluation Measure (ITEM) is a multidomain questionnaire assessing immersion, motivation, cognitive load, usability, and debriefing. Although cognitive interviewing informed its original development, less is known about how ITEM questions function when used in a different linguistic and educational context. Objective This study aimed to evaluate the clarity, comprehensibility, and response processes of ITEM among native Arabic-speaking medical students enrolled in an English-medium medical program. We also sought to identify linguistic, referential, and contextual sources of comprehension difficulty and use these findings to inform proposed item-level refinements. Methods We conducted a qualitative cognitive interviewing study with 10 third-year and fourth-year medical students at United Arab Emirates University following VR-based clinical skills activities. Using concurrent think-aloud and verbal probing techniques, participants explained how they interpreted each ITEM question and arrived at their responses. Interview transcripts, recordings, and interviewer notes were independently reviewed using a descriptive, item-focused approach. Participant feedback was examined for recurring comprehension problems, and proposed item-level decisions were reviewed through research team consensus. Item-level saturation was reached after 8 interviews and confirmed with 2 additional interviews. Results Participants identified comprehension or response-process problems in 47.5% (19/40) of ITEM questions. Difficulties occurred across all 5 domains and were most frequent in immersion (6/9, 66.7%) and usability (6/10, 60%), followed by debriefing (4/5, 80%), motivation (2/10, 20%), and cognitive load (1/6, 16.7%). Common difficulties involved nonspecific references such as “activity” and “technology,” unfamiliar or abstract terminology, negative wording, and unclear temporal or contextual framing. For example, “concern” was interpreted by some participants as worry rather than focus, and “mentally demanding” was interpreted in relation to mental health rather than cognitive effort. Cognitive interviewing informed different item-level decisions rather than a uniform revision: 11 problematic questions received proposed wording or contextual revisions, 2 received presentation-only modifications, and 6 were retained unchanged where clarification could introduce a meaning not clearly established in the original item. Six additional questions received limited contextual or terminology-standardizing changes for consistency. No items were removed, and the original 5-domain structure was retained. Conclusions Cognitive interviewing identified response-process difficulties that were not apparent from the questionnaire wording alone and provided a systematic basis for determining when clarification was warranted and when the original wording should be preserved. The findings extend response-process evidence for ITEM and illustrate the value of examining established educational measures in new linguistic and educational settings. The proposed refinements provide a foundation for further cognitive testing and psychometric evaluation across immersive learning contexts.

Authors

Publication Details

Journal
JMIR Medical Education
Published
2026-09-28
DOI
https://doi.org/10.2196/95904
Primary Topic
Virtual Reality Applications and Impacts
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Refining the Immersive Technology Evaluation Measure for Virtual Reality Clinical Skills Among Native Arabic-Speaking Medical Students: Qualitative Study

Taleb Mohamed Almansoori, Afaf Sulaiman Alblooshi, Faten A. Alradini, Alexander Kieu et al.
JMIR Medical Education
Virtual Reality Applications and Impacts
article

Refining the Immersive Technology Evaluation Measure for Virtual Reality Clinical Skills Among Native Arabic-Speaking Medical Students: Qualitative Study

Taleb Mohamed Almansoori, Afaf Sulaiman Alblooshi, Faten A. Alradini, Alexander Kieu, Falah Mohammed AlMarzooqi, Anmol Punn, Marwa Gaffar Alameen
article en

Abstract

Abstract Background Extended reality technologies, including virtual reality (VR), augmented reality, and mixed reality, are increasingly used in medical education to create immersive and interactive learning environments. As these modalities expand, validated instruments are needed to measure learners’ experiences accurately. The Immersive Technology Evaluation Measure (ITEM) is a multidomain questionnaire assessing immersion, motivation, cognitive load, usability, and debriefing. Although cognitive interviewing informed its original development, less is known about how ITEM questions function when used in a different linguistic and educational context. Objective This study aimed to evaluate the clarity, comprehensibility, and response processes of ITEM among native Arabic-speaking medical students enrolled in an English-medium medical program. We also sought to identify linguistic, referential, and contextual sources of comprehension difficulty and use these findings to inform proposed item-level refinements. Methods We conducted a qualitative cognitive interviewing study with 10 third-year and fourth-year medical students at United Arab Emirates University following VR-based clinical skills activities. Using concurrent think-aloud and verbal probing techniques, participants explained how they interpreted each ITEM question and arrived at their responses. Interview transcripts, recordings, and interviewer notes were independently reviewed using a descriptive, item-focused approach. Participant feedback was examined for recurring comprehension problems, and proposed item-level decisions were reviewed through research team consensus. Item-level saturation was reached after 8 interviews and confirmed with 2 additional interviews. Results Participants identified comprehension or response-process problems in 47.5% (19/40) of ITEM questions. Difficulties occurred across all 5 domains and were most frequent in immersion (6/9, 66.7%) and usability (6/10, 60%), followed by debriefing (4/5, 80%), motivation (2/10, 20%), and cognitive load (1/6, 16.7%). Common difficulties involved nonspecific references such as “activity” and “technology,” unfamiliar or abstract terminology, negative wording, and unclear temporal or contextual framing. For example, “concern” was interpreted by some participants as worry rather than focus, and “mentally demanding” was interpreted in relation to mental health rather than cognitive effort. Cognitive interviewing informed different item-level decisions rather than a uniform revision: 11 problematic questions received proposed wording or contextual revisions, 2 received presentation-only modifications, and 6 were retained unchanged where clarification could introduce a meaning not clearly established in the original item. Six additional questions received limited contextual or terminology-standardizing changes for consistency. No items were removed, and the original 5-domain structure was retained. Conclusions Cognitive interviewing identified response-process difficulties that were not apparent from the questionnaire wording alone and provided a systematic basis for determining when clarification was warranted and when the original wording should be preserved. The findings extend response-process evidence for ITEM and illustrate the value of examining established educational measures in new linguistic and educational settings. The proposed refinements provide a foundation for further cognitive testing and psychometric evaluation across immersive learning contexts.

JMIR Medical EducationVol. 12
Quality Education
Openalex Percentile: Top 9%
Virtual Reality Applications and Impacts
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.