AI powered handwritten answer script evaluation framework for automated assessment of text and figure based responses

Manual evaluation of handwritten answer scripts is time-consuming and subjective, and grows increasingly impractical as enrolment rises; moreover, most existing automated grading systems support only typed text and cannot assess figure-based responses, leaving a substantial gap in educational technology. To address this problem, this research presents a computer-based evaluation framework for handwritten answer scripts containing both textual and graphical responses. Scripts are digitized through optical character recognition (OCR), and the recognized text is evaluated using four complementary techniques: a writing-quality score, a keyword matching score, a similarity score, and a reference-overlap score that quantifies verbatim copying from the expected answer. Two text evaluation cases are compared on a validation set of 25 student responses: the first case employs TF-IDF-based cosine similarity and reaches 83.8% agreement with examiner-assigned marks, while the second replaces it with NLP-based semantic similarity using spaCy word embeddings, improving agreement to 87.4%; this improvement is statistically significant (Wilcoxon signed-rank test, p = 0.007). Separately, figure evaluation compares student-drawn diagrams with reference figures using deep convolutional (MobileNetV2) feature embeddings matched by cosine similarity, achieving 85.1% agreement, versus 81.1% for a Hu-moment-based Euclidean distance approach. Agreement with human examiners is further quantified through MAE, RMSE, Pearson and Spearman correlation, and quadratic weighted kappa. These results imply that combining semantic text similarity with deep feature-based figure comparison can provide fair and consistent grading support for educators; broader deployment in universities and standardized testing will require validation on larger, more diverse datasets.

Authors

Institutions

Publication Details

Journal
Discover Artificial Intelligence
Published
2026-09-25
DOI
https://doi.org/10.1007/s44163-026-02181-4
Primary Topic
Handwritten Text Recognition Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

AI powered handwritten answer script evaluation framework for automated assessment of text and figure based responses

Ishtiaq Ahammad, Md. Ashikur Rahman Khan, Md Nurul Amin
Discover Artificial Intelligence
Handwritten Text Recognition Techniques
article

AI powered handwritten answer script evaluation framework for automated assessment of text and figure based responses

Ishtiaq Ahammad, Md. Ashikur Rahman Khan, Md Nurul Amin
article en

Abstract

Manual evaluation of handwritten answer scripts is time-consuming and subjective, and grows increasingly impractical as enrolment rises; moreover, most existing automated grading systems support only typed text and cannot assess figure-based responses, leaving a substantial gap in educational technology. To address this problem, this research presents a computer-based evaluation framework for handwritten answer scripts containing both textual and graphical responses. Scripts are digitized through optical character recognition (OCR), and the recognized text is evaluated using four complementary techniques: a writing-quality score, a keyword matching score, a similarity score, and a reference-overlap score that quantifies verbatim copying from the expected answer. Two text evaluation cases are compared on a validation set of 25 student responses: the first case employs TF-IDF-based cosine similarity and reaches 83.8% agreement with examiner-assigned marks, while the second replaces it with NLP-based semantic similarity using spaCy word embeddings, improving agreement to 87.4%; this improvement is statistically significant (Wilcoxon signed-rank test, p = 0.007). Separately, figure evaluation compares student-drawn diagrams with reference figures using deep convolutional (MobileNetV2) feature embeddings matched by cosine similarity, achieving 85.1% agreement, versus 81.1% for a Hu-moment-based Euclidean distance approach. Agreement with human examiners is further quantified through MAE, RMSE, Pearson and Spearman correlation, and quadratic weighted kappa. These results imply that combining semantic text similarity with deep feature-based figure comparison can provide fair and consistent grading support for educators; broader deployment in universities and standardized testing will require validation on larger, more diverse datasets.

Discover Artificial IntelligenceVol. 6(1)
Noakhali Science and Technology University (BD)
Quality Education
Openalex Percentile: Top 14%
Handwritten Text Recognition Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.