Using large language models to analyze student reflections on generative AI in graduate coursework

This exploratory study examines graduate students’ perceptions of generative AI use in coursework and evaluates the feasibility of using large language models (LLMs) to support thematic coding of brief open-ended responses. Students ( n = 20) enrolled in a systems analysis and design course completed a post-course survey consisting of Likert-scale items and eight open-ended questions, six retrospective and two prospective or hypothetical. Quantitative responses were summarized descriptively, while qualitative responses were independently coded by two human researchers and two LLMs (GPT-4o and Claude 3.7 Sonnet) using a shared, human-developed codebook. Inter-rater reliability was assessed using Cohen’s and Fleiss’ kappa. Students reported high levels of engagement, motivation, perceived learning support, and perceived efficiency when using course-specific LLM applications; however, the study did not assess objective learning outcomes. Human-human agreement ranged from moderate to almost perfect (κ = 0.52–0.84), whereas average human-LLM agreement was lower and more variable (κ = 0.36–0.62). Claude aligned more closely with the human coders than GPT-4o across all eight questions when the two human-model kappas were averaged. Agreement was highest for concrete, task-focused prompts and lower for several reflective and prospective prompts, where the LLMs missed human-identified themes, fragmented broad themes into narrower subcodes, or assigned substantive meaning to non-responses. Hybrid intelligent feedback is used as a retrospective sensitizing framework rather than as a prospectively operationalized intervention. The findings support a bounded role for LLMs in qualitative educational research: models can assist with screening, discrepancy detection, and code suggestions, but human researchers remain necessary for contextual interpretation, validation, and adjudication.

Authors

Institutions

Publication Details

Journal
Discover Education
Published
2026-09-29
DOI
https://doi.org/10.1007/s44217-026-02231-0
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Using large language models to analyze student reflections on generative AI in graduate coursework

Billie S. Anderson, Tyler Price, Viktors J. Muiznieks
Discover Education
Artificial Intelligence in Healthcare and Education
article

Using large language models to analyze student reflections on generative AI in graduate coursework

Billie S. Anderson, Tyler Price, Viktors J. Muiznieks
article en

Abstract

This exploratory study examines graduate students’ perceptions of generative AI use in coursework and evaluates the feasibility of using large language models (LLMs) to support thematic coding of brief open-ended responses. Students ( n = 20) enrolled in a systems analysis and design course completed a post-course survey consisting of Likert-scale items and eight open-ended questions, six retrospective and two prospective or hypothetical. Quantitative responses were summarized descriptively, while qualitative responses were independently coded by two human researchers and two LLMs (GPT-4o and Claude 3.7 Sonnet) using a shared, human-developed codebook. Inter-rater reliability was assessed using Cohen’s and Fleiss’ kappa. Students reported high levels of engagement, motivation, perceived learning support, and perceived efficiency when using course-specific LLM applications; however, the study did not assess objective learning outcomes. Human-human agreement ranged from moderate to almost perfect (κ = 0.52–0.84), whereas average human-LLM agreement was lower and more variable (κ = 0.36–0.62). Claude aligned more closely with the human coders than GPT-4o across all eight questions when the two human-model kappas were averaged. Agreement was highest for concrete, task-focused prompts and lower for several reflective and prospective prompts, where the LLMs missed human-identified themes, fragmented broad themes into narrower subcodes, or assigned substantive meaning to non-responses. Hybrid intelligent feedback is used as a retrospective sensitizing framework rather than as a prospectively operationalized intervention. The findings support a bounded role for LLMs in qualitative educational research: models can assist with screening, discrepancy detection, and code suggestions, but human researchers remain necessary for contextual interpretation, validation, and adjudication.

Discover EducationVol. 5(1)
Southern New Hampshire University (US)
Quality Education
Openalex Percentile: Top 16%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Using large language models to analyze student reflections on generative AI in graduate coursework — Billie S. Anderson, Tyler Price, et al. · Discover Education (2026) | TGRS Research Map | TGRS