Optimizing AI-based assessment in history education: The impact of prompt engineering on scoring performance in a morphologically rich language

This study investigates the effectiveness and prompt sensitivity of leading large language model (LLM)-based AI tools (ChatGPT, Gemini, Claude, Deepseek) in grading open-ended exam questions in an undergraduate history course (Atatürk’s Principles and History of Revolution), compared to human evaluators. A comprehensive dataset comprising 72 distinct open-ended responses (collected from 24 students) was scored by both the course instructor and seven AI models using five different prompts with varying levels of detail. These prompts ranged from a basic zero-shot instruction to progressively more structured designs: expected topic headings, a weighted criterion-based rubric, partial coverage with language proficiency, and relative (norm-referenced) scoring. Correlation and statistical analyses (paired samples t-test, Wilcoxon signed-rank test) revealed that while models showed low agreement with the human evaluator when using unstructured, basic prompts (zero-shot), they achieved high agreement (r > .80) when provided with structured prompts and explicit rubrics, particularly in the cases of Gemini and Claude. The findings highlight that for effective AI-based assessment, prompt design and the definition of criteria are more critical than model selection in mitigating "generosity bias." Generosity bias here denotes the models’ systematic tendency to assign higher scores than the human rater; the qualitative evidence indicates that it stems primarily from the models rewarding fluent, lengthy, and well-structured answers even when their factual content is incomplete, a tendency that explicit rubrics substantially reduced. Designed as an exploratory case study, these results provide significant empirical evidence regarding the optimization of AI as an assistive assessment tool, specifically within the context of Turkish, a morphologically rich language.

Authors

Institutions

Publication Details

Journal
Journal of Educational Technology and Online Learning
Published
2026-09-30
DOI
https://doi.org/10.31681/jetol.1882892
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Optimizing AI-based assessment in history education: The impact of prompt engineering on scoring performance in a morphologically rich language

Emine AŞÇI, Yunus Özdemir
Journal of Educational Technology and Online Learning
Artificial Intelligence in Healthcare and Education
article

Optimizing AI-based assessment in history education: The impact of prompt engineering on scoring performance in a morphologically rich language

Emine AŞÇI, Yunus Özdemir
article en

Abstract

This study investigates the effectiveness and prompt sensitivity of leading large language model (LLM)-based AI tools (ChatGPT, Gemini, Claude, Deepseek) in grading open-ended exam questions in an undergraduate history course (Atatürk’s Principles and History of Revolution), compared to human evaluators. A comprehensive dataset comprising 72 distinct open-ended responses (collected from 24 students) was scored by both the course instructor and seven AI models using five different prompts with varying levels of detail. These prompts ranged from a basic zero-shot instruction to progressively more structured designs: expected topic headings, a weighted criterion-based rubric, partial coverage with language proficiency, and relative (norm-referenced) scoring. Correlation and statistical analyses (paired samples t-test, Wilcoxon signed-rank test) revealed that while models showed low agreement with the human evaluator when using unstructured, basic prompts (zero-shot), they achieved high agreement (r > .80) when provided with structured prompts and explicit rubrics, particularly in the cases of Gemini and Claude. The findings highlight that for effective AI-based assessment, prompt design and the definition of criteria are more critical than model selection in mitigating "generosity bias." Generosity bias here denotes the models’ systematic tendency to assign higher scores than the human rater; the qualitative evidence indicates that it stems primarily from the models rewarding fluent, lengthy, and well-structured answers even when their factual content is incomplete, a tendency that explicit rubrics substantially reduced. Designed as an exploratory case study, these results provide significant empirical evidence regarding the optimization of AI as an assistive assessment tool, specifically within the context of Turkish, a morphologically rich language.

Journal of Educational Technology and Online LearningVol. 9(3)
Bolu Abant İzzet Baysal University (TR)
Quality Education
Openalex Percentile: Top 15%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Optimizing AI-based assessment in history education: The impact of prompt engineering on scoring performance in a morphologically rich language — Emine AŞÇI, Yunus Özdemir · Journal of Educational Technology and Online Learning (2026) | TGRS Research Map | TGRS