Analyzing Educational Video Transcripts with LLMs: Accuracy, Named Entities, and Interpretability
Digital educational media platforms such as YouTube are increasingly used by teachers and learners, creating a growing record of instructional discourse. For educational data mining (EDM) and learning analytics researchers, transcript-derived text offers a scalable entry point for corpus-level analysis, but automatically generated transcripts can contain errors that disproportionately affect named entities and downstream applications. We report an initial validation study of a transcript-to-entity workflow for educational video analysis. The workflow benchmarks multiple automatic speech recognition (ASR) systems, evaluates GPT-4o and GPT-5.5 against a 600-token human consensus named-entity benchmark, compares both LLMs with an off-the-shelf spaCy transformer NER pipeline, and applies the model best aligned with the study-specific coding frame to a 48-episode corpus from the Crash Course U.S. History series. The results reveal an important task-dependent trade-off: spaCy performed best at general entity detection, whereas GPT-5.5 showed stronger aggregate performance for codebook-constrained entity typing on gold-entity rows. This finding supports the use of GPT-5.5 for the study's type-specific corpus profiling without implying universal model superiority. Across the full corpus, PERSON, GPE, and NORP were the most frequent model-coded categories. Topic modeling comparisons further show that transcript-conditioning choices can change thematic separation and interpretability. Together, the findings demonstrate why transcript quality and entity validity should be evaluated according to the downstream educational research task. The workflow provides a reproducible starting point whose corpus-scale outputs can support navigation, sampling, and targeted human review rather than definitive annotation. Codes and data are at https://github.com/wwang93/JEDM-Paper-Pipeline.git.
Authors
- Cody Pritchard (ORCID: https://orcid.org/0009-0006-9080-8573)
- Joshua Rosenberg
- Wei Wang
Institutions
- University of Tennessee at Knoxville (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-10
- DOI
- https://doi.org/10.5281/zenodo.22689986
- Primary Topic
- Online Learning and Analytics
- Type
- article
- Field-Weighted Citation Impact
- 0.00