Analysis of Student Performance Prediction Using the Bi‐LSTM Technique Focused on the Indian Higher Education Context Using a Synthetic Dataset

ABSTRACT Student performance prediction is a critical task for improving educational outcomes and enabling data‐driven decision‐making in higher education. This study proposes a context‐specific predictive framework for Indian higher education that integrates AISHE institutional indicators with student‐level behavioral data through multi‐level data fusion, incorporates SHAP‐based interpretability to support policy and intervention decisions, and employs a Bi‐LSTM + Nadam architecture to achieve strong predictive performance. The framework integrates multi‐level heterogeneous data, combining institutional indicators from AISHE with student‐level academic, behavioral, and demographic features. To address the lack of unified large‐scale datasets, a synthetic data generation approach is employed to simulate multi‐institutional environments with temporal student records, supported by a mapping mechanism aligning student and institutional attributes. Experimental results demonstrate that the proposed model achieves 93.1% accuracy, significantly outperforming traditional machine learning approaches by approximately 13%. The model maintains precision of 93.0%, recall of 93.1%, and F1‐scores of 93% across all performance categories, confirming robust and balanced classification. Convergence analysis shows that Nadam enables faster and more stable training, while the Bi‐LSTM architecture improves performance by 1.9% over unidirectional LSTM by capturing bidirectional temporal dependencies. Feature importance analysis reveals that academic and behavioral factors dominate prediction, with GPA, attendance, and assignment submission contributing approximately 42.7% of the total normalized SHAP importance, while demographic factors have minimal influence, ensuring reduced bias. ROC‐AUC analysis confirms excellent discrimination across all categories, with a macro‐average AUC of 0.993 (range: 0.989–0.997), fully consistent with the reported accuracy and recall metrics. The proposed framework is suitable for early warning systems, targeted interventions, and academic decision support, making it highly applicable in real‐world educational settings.

Authors

Institutions

Publication Details

Journal
Engineering Reports
Published
2026-09-30
DOI
https://doi.org/10.1002/eng2.71084
Primary Topic
Online Learning and Analytics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Analysis of Student Performance Prediction Using the Bi‐LSTM Technique Focused on the Indian Higher Education Context Using a Synthetic Dataset

Jemal Abate, Kalkundri Ravi, Sudhindra B. Deshpande, Praveen Kalkundri et al.
Engineering Reports
Online Learning and Analytics
article

Analysis of Student Performance Prediction Using the Bi‐LSTM Technique Focused on the Indian Higher Education Context Using a Synthetic Dataset

Jemal Abate, Kalkundri Ravi, Sudhindra B. Deshpande, Praveen Kalkundri, Pratijnya Ajawan, Raghavendra Jadhav
article en

Abstract

ABSTRACT Student performance prediction is a critical task for improving educational outcomes and enabling data‐driven decision‐making in higher education. This study proposes a context‐specific predictive framework for Indian higher education that integrates AISHE institutional indicators with student‐level behavioral data through multi‐level data fusion, incorporates SHAP‐based interpretability to support policy and intervention decisions, and employs a Bi‐LSTM + Nadam architecture to achieve strong predictive performance. The framework integrates multi‐level heterogeneous data, combining institutional indicators from AISHE with student‐level academic, behavioral, and demographic features. To address the lack of unified large‐scale datasets, a synthetic data generation approach is employed to simulate multi‐institutional environments with temporal student records, supported by a mapping mechanism aligning student and institutional attributes. Experimental results demonstrate that the proposed model achieves 93.1% accuracy, significantly outperforming traditional machine learning approaches by approximately 13%. The model maintains precision of 93.0%, recall of 93.1%, and F1‐scores of 93% across all performance categories, confirming robust and balanced classification. Convergence analysis shows that Nadam enables faster and more stable training, while the Bi‐LSTM architecture improves performance by 1.9% over unidirectional LSTM by capturing bidirectional temporal dependencies. Feature importance analysis reveals that academic and behavioral factors dominate prediction, with GPA, attendance, and assignment submission contributing approximately 42.7% of the total normalized SHAP importance, while demographic factors have minimal influence, ensuring reduced bias. ROC‐AUC analysis confirms excellent discrimination across all categories, with a macro‐average AUC of 0.993 (range: 0.989–0.997), fully consistent with the reported accuracy and recall metrics. The proposed framework is suitable for early warning systems, targeted interventions, and academic decision support, making it highly applicable in real‐world educational settings.

Engineering ReportsVol. 8(10)
Haramaya University (ET), KLS Gogte Institute of Technology
Peace, Justice and strong institutions
Openalex Percentile: Top 6%
Online Learning and Analytics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.