Comparative Analysis of Base Machine-Learning Models for Dyslexia Prediction Using Behavioral Learning Data

Dyslexia is a specific learning disorder that affects reading, decoding, spelling and fluent word recognition. Early identification can facilitate timely educational intervention. This study comparatively evaluates three machine-learning (ML) base classifiers—Random Forest (RF), XGBoost and Extra Trees (ET)—for dyslexia prediction using behavioural learning data. The supplied Dyt-desktop dataset contains 3,644 observations, 196 predictors and 392 dyslexia-positive cases. Stratified five-fold cross-validation was applied with class-weighted classifiers. Performance was assessed using accuracy, precision, recall, specificity, F1-score, ROC-AUC and Matthews Correlation Coefficient (MCC). XGBoost achieved the strongest overall base-model performance, with 90.64% accuracy, 64.41% precision, 29.08% recall, 98.06% specificity, 40.07% F1-score, 0.8845 ROC-AUC and 0.3912 MCC. RF-RFE + XGBoost improved performance to 90.81% accuracy and 31.89% recall. The previously developed probability-level ensemble achieved 71.17% recall and 0.4644 MCC, although its accuracy decreased to 85.70%. The findings demonstrate that accuracy alone is inadequate for evaluating dyslexia prediction under substantial class imbalance. XGBoost is the strongest individual base learner, while ensemble learning provides greater sensitivity when screening is prioritized.

Authors

Institutions

Publication Details

Journal
Iconic Research and Engineering Journals
Published
2026-09-07
DOI
https://doi.org/10.64388/irev10i3-1722828
Primary Topic
Reading and Literacy Development
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Comparative Analysis of Base Machine-Learning Models for Dyslexia Prediction Using Behavioral Learning Data

Faruk Umar Ambursa, Femi Adeluyi, Chigozirim Ajaegu, Patrick Ufomba Nwogu
Iconic Research and Engineering Journals
Reading and Literacy Development
article

Comparative Analysis of Base Machine-Learning Models for Dyslexia Prediction Using Behavioral Learning Data

Faruk Umar Ambursa, Femi Adeluyi, Chigozirim Ajaegu, Patrick Ufomba Nwogu
article en

Abstract

Dyslexia is a specific learning disorder that affects reading, decoding, spelling and fluent word recognition. Early identification can facilitate timely educational intervention. This study comparatively evaluates three machine-learning (ML) base classifiers—Random Forest (RF), XGBoost and Extra Trees (ET)—for dyslexia prediction using behavioural learning data. The supplied Dyt-desktop dataset contains 3,644 observations, 196 predictors and 392 dyslexia-positive cases. Stratified five-fold cross-validation was applied with class-weighted classifiers. Performance was assessed using accuracy, precision, recall, specificity, F1-score, ROC-AUC and Matthews Correlation Coefficient (MCC). XGBoost achieved the strongest overall base-model performance, with 90.64% accuracy, 64.41% precision, 29.08% recall, 98.06% specificity, 40.07% F1-score, 0.8845 ROC-AUC and 0.3912 MCC. RF-RFE + XGBoost improved performance to 90.81% accuracy and 31.89% recall. The previously developed probability-level ensemble achieved 71.17% recall and 0.4644 MCC, although its accuracy decreased to 85.70%. The findings demonstrate that accuracy alone is inadequate for evaluating dyslexia prediction under substantial class imbalance. XGBoost is the strongest individual base learner, while ensemble learning provides greater sensitivity when screening is prioritized.

Iconic Research and Engineering JournalsVol. 10(3)
National Open University of Nigeria (NG)
Quality Education
Openalex Percentile: Top 5%
Reading and Literacy Development
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.