Comparative Analysis of K-Means and Random Forest for Student Academic Performance Prediction

The increasing availability of educational data has created opportunities for machine learning (ML) techniques to support academic performance analysis, early identification of students at risk of poor outcomes, and evidence-based educational decision-making. This study presents a comparative analysis of K-Means clustering and Random Forest (RF) for student academic performance prediction. Although both techniques can discover useful patterns in educational datasets, they differ fundamentally in their learning paradigms: K-Means is an unsupervised clustering algorithm that identifies groups of students with similar characteristics, whereas Random Forest is a supervised ensemble-learning algorithm capable of directly predicting predefined academic outcomes. The study adopts the UCI Student Performance dataset, which contains demographic, social, school-related, behavioural, and academic variables. The proposed methodology involves data preprocessing, feature transformation, exploratory analysis, K-Means clustering, Random Forest classification, and comparative evaluation using accuracy, precision, recall, F1-score, cluster quality, and computational considerations. Recent studies indicate that tree-based methods, particularly Random Forest and related ensemble models, remain widely used in student-performance prediction, while clustering provides useful insights into heterogeneous student profiles. The analysis demonstrates that the two algorithms should not be regarded as direct substitutes: K-Means is primarily valuable for discovering latent student groups, whereas Random Forest is more appropriate when the objective is explicit prediction of academic-performance categories. The study therefore highlights the complementary rather than identical roles of clustering and supervised learning in educational data mining.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-03
DOI
https://doi.org/10.5281/zenodo.23121737
Primary Topic
Online Learning and Analytics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Comparative Analysis of K-Means and Random Forest for Student Academic Performance Prediction

Mustapha Malami Idina, Mubarak Jibril Yeldu
Zenodo (CERN European Organization for Nuclear Research)
Online Learning and Analytics
article

Comparative Analysis of K-Means and Random Forest for Student Academic Performance Prediction

Mustapha Malami Idina, Mubarak Jibril Yeldu
article en

Abstract

The increasing availability of educational data has created opportunities for machine learning (ML) techniques to support academic performance analysis, early identification of students at risk of poor outcomes, and evidence-based educational decision-making. This study presents a comparative analysis of K-Means clustering and Random Forest (RF) for student academic performance prediction. Although both techniques can discover useful patterns in educational datasets, they differ fundamentally in their learning paradigms: K-Means is an unsupervised clustering algorithm that identifies groups of students with similar characteristics, whereas Random Forest is a supervised ensemble-learning algorithm capable of directly predicting predefined academic outcomes. The study adopts the UCI Student Performance dataset, which contains demographic, social, school-related, behavioural, and academic variables. The proposed methodology involves data preprocessing, feature transformation, exploratory analysis, K-Means clustering, Random Forest classification, and comparative evaluation using accuracy, precision, recall, F1-score, cluster quality, and computational considerations. Recent studies indicate that tree-based methods, particularly Random Forest and related ensemble models, remain widely used in student-performance prediction, while clustering provides useful insights into heterogeneous student profiles. The analysis demonstrates that the two algorithms should not be regarded as direct substitutes: K-Means is primarily valuable for discovering latent student groups, whereas Random Forest is more appropriate when the objective is explicit prediction of academic-performance categories. The study therefore highlights the complementary rather than identical roles of clustering and supervised learning in educational data mining.

Zenodo (CERN European Organization for Nuclear Research)
Openalex Percentile: Top 6%
Online Learning and Analytics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Comparative Analysis of K-Means and Random Forest for Student Academic Performance Prediction — Mustapha Malami Idina, Mubarak Jibril Yeldu · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS