Evaluating Gemma 3 27B for local undergraduate physics tutoring

Abstract Physics education research has developed a range of student-centered approaches that emphasize conceptual understanding and problem solving, creating interest in whether large language models (LLMs) can provide complementary, on-demand tutoring support. Locally deployable models are particularly relevant for institutions seeking greater control over data handling and system operation, but their capability relative to commercial models and their reliability in authentic learning interactions remain insufficiently characterized. We introduce mlphys101, a multilingual multiple-choice benchmark of introductory physics questions, and evaluate Gemma 3 27B on its German variant in three quantization configurations. The locally deployed model achieves high benchmark accuracy, although performance is lower and more variable on conceptual and multi-step reasoning tasks than that of a commercial cloud baseline. We subsequently deploy the same model as a tutoring chatbot in a classroom field experiment with first-semester engineering students. Survey responses indicate that explanations were generally comprehensible, but the practical usefulness of the system was limited by frequent prompt reformulation, occasional incorrect or incomplete responses, difficulties in using information from the ongoing interaction, and response latency. These findings demonstrate that strong performance on a controlled physics benchmark does not necessarily translate into reliable tutoring behavior. Locally deployed LLMs may therefore provide useful support for selected learning activities, but their suitability for educational use should be evaluated not only through domain benchmarks but also under realistic, multi-turn interaction conditions.

Authors

Institutions

Publication Details

Journal
Discover Artificial Intelligence
Published
2026-10-07
DOI
https://doi.org/10.1007/s44163-026-02296-8
Primary Topic
Intelligent Tutoring Systems and Adaptive Learning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Evaluating Gemma 3 27B for local undergraduate physics tutoring

Marcel Völschow, P. Buczek, Philipp Riefer, Peter Steinbach et al.
Discover Artificial Intelligence
Intelligent Tutoring Systems and Adaptive Learning
article

Evaluating Gemma 3 27B for local undergraduate physics tutoring

Marcel Völschow, P. Buczek, Philipp Riefer, Peter Steinbach, Adelina Jasko
article en

Abstract

Abstract Physics education research has developed a range of student-centered approaches that emphasize conceptual understanding and problem solving, creating interest in whether large language models (LLMs) can provide complementary, on-demand tutoring support. Locally deployable models are particularly relevant for institutions seeking greater control over data handling and system operation, but their capability relative to commercial models and their reliability in authentic learning interactions remain insufficiently characterized. We introduce mlphys101, a multilingual multiple-choice benchmark of introductory physics questions, and evaluate Gemma 3 27B on its German variant in three quantization configurations. The locally deployed model achieves high benchmark accuracy, although performance is lower and more variable on conceptual and multi-step reasoning tasks than that of a commercial cloud baseline. We subsequently deploy the same model as a tutoring chatbot in a classroom field experiment with first-semester engineering students. Survey responses indicate that explanations were generally comprehensible, but the practical usefulness of the system was limited by frequent prompt reformulation, occasional incorrect or incomplete responses, difficulties in using information from the ongoing interaction, and response latency. These findings demonstrate that strong performance on a controlled physics benchmark does not necessarily translate into reliable tutoring behavior. Locally deployed LLMs may therefore provide useful support for selected learning activities, but their suitability for educational use should be evaluated not only through domain benchmarks but also under realistic, multi-turn interaction conditions.

Discover Artificial IntelligenceVol. 6(1)
Helmholtz-Zentrum Dresden-Rossendorf (DE), Deutsches Elektronen-Synchrotron DESY (DE), HAW Hamburg (DE)
Openalex Percentile: Top 12%
Intelligent Tutoring Systems and Adaptive Learning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Evaluating Gemma 3 27B for local undergraduate physics tutoring — Marcel Völschow, P. Buczek, et al. · Discover Artificial Intelligence (2026) | TGRS Research Map | TGRS