Uncovering the drivers of IPO underpricing in Malaysia through machine learning benchmarking

Abstract Purpose In this study, IPO underpricing in a fixed-price market is examined through two distinct objectives: (1) a predictive objective identifying the machine learning model that best generalizes out-of-sample, and (2) an economic interpretation objective using model-agnostic interpretability tools (SHAP) and formal statistical tests (bootstrapped hierarchical regression) to assess which economic mechanisms are consistent with the observed predictive patterns. Conventional linear models with additive feature effects are inefficient for understanding moderation effects involving interactions. This is done through two methodological innovations: (1) explicit non-linear feature engineering (interaction terms to describe moderation effects), allowing linear models to learn complex relationships, and (2) model-agnostic interpretability of tree-based ensembles through SHAP analysis to numerically quantify interaction strengths. It combines the predictive stability of regularized linear models (SVR-Linear: R 2 = 0.512) with the interpretability of non-linear models (Random Forest SHAP) for hypothesis testing. Design/methodology The data employed in this paper include 350 Malaysian IPOs of the respective years 2004–2021. It uses seven machine learning algorithms: Support Vector Regression (SVR), Random Forest, XGBoost, LightGBM, k-Nearest Neighbors, Multilayer Perceptron, and Linear Regression. The evaluated algorithms are systematically assessed using 5-fold cross-validation. In terms of feature importance, SHAP (Shapley Additive exPlanations) values are used to quantify feature importance and characterize associations regarding investor demand, offer price, and listing board. Formal hypothesis testing is conducted via bootstrapped hierarchical regression. Findings SVR-Linear offers better predictive performance (test R 2 = 0.512, CV R 2 = 0.486 ± 0.087), whereas ensemble techniques are prone to overfitting. SHAP analysis shows that the over-subscription ratio (OSR) contributes 40.7%, followed by a knowledge gap (15.6%). Hierarchical regression is used to establish strong moderating roles: investor demand is a significant moderator of the knowledge gap-underpricing relationship (β = 0.0284, p < 0.01), and intense demand (OSR > 30x) dilutes the effect of information asymmetry by 73%. There is marginal moderation in the offer price (β = 0.0198, p = 0.073) and no significant effect by listing board differences (β = 0.0089, p = 0.496). The threshold analysis shows that when investor demand is high, cascading demand effects lead to fundamental valuation concerns and patterns consistent with the predictions of Welch’s (1992) bandwagon theory. Practical implications The findings from this research can benefit a variety of stakeholders. The stockholders can benefit from this research and minimize underpricing. Furthermore, regulatory bodies will consider book-building processes for large IPOs when there is a signal of demand. Additionally, the institutional investors should avoid oversubscription. Furthermore, it can be concluded that precise targeting policies can be achieved by measuring interaction effects. Originality/value The study contributes in three ways, including (1) the first application of SHAP-based interpretability to IPO underpricing, which allows the identification of non-linear moderation effects that are not observed by conventional regression, (2) a rigorous overfitting diagnostic can be evaluated using hierarchical regression with bootstrapped confidence intervals, a feature that identifies the gap between black-box prediction and inferential rigor; and (3) formal statistical validation of SHAP results through bootstrapped hierarchical regression, a property that has not been previously available in financial machine learning studies. Explainable AI, coupled with causal inference, offers a transparent and replicable analytical framework for financial research that requires predictive validity and theoretical elucidation.

Authors

Publication Details

Journal
Discover Artificial Intelligence
Published
2026-09-30
DOI
https://doi.org/10.1007/s44163-026-02390-x
Primary Topic
Financial Distress and Bankruptcy Prediction
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Uncovering the drivers of IPO underpricing in Malaysia through machine learning benchmarking

Boon Heng Teh, Khalid Al Qatiti, Ali Albada, Rabie Α. Ramadan et al.
Discover Artificial Intelligence
Financial Distress and Bankruptcy Prediction
article

Uncovering the drivers of IPO underpricing in Malaysia through machine learning benchmarking

Boon Heng Teh, Khalid Al Qatiti, Ali Albada, Rabie Α. Ramadan, Soo-Wah Low, Eimad Eldin Abusham, Chui Zi Ong
article en

Abstract

Abstract Purpose In this study, IPO underpricing in a fixed-price market is examined through two distinct objectives: (1) a predictive objective identifying the machine learning model that best generalizes out-of-sample, and (2) an economic interpretation objective using model-agnostic interpretability tools (SHAP) and formal statistical tests (bootstrapped hierarchical regression) to assess which economic mechanisms are consistent with the observed predictive patterns. Conventional linear models with additive feature effects are inefficient for understanding moderation effects involving interactions. This is done through two methodological innovations: (1) explicit non-linear feature engineering (interaction terms to describe moderation effects), allowing linear models to learn complex relationships, and (2) model-agnostic interpretability of tree-based ensembles through SHAP analysis to numerically quantify interaction strengths. It combines the predictive stability of regularized linear models (SVR-Linear: R 2 = 0.512) with the interpretability of non-linear models (Random Forest SHAP) for hypothesis testing. Design/methodology The data employed in this paper include 350 Malaysian IPOs of the respective years 2004–2021. It uses seven machine learning algorithms: Support Vector Regression (SVR), Random Forest, XGBoost, LightGBM, k-Nearest Neighbors, Multilayer Perceptron, and Linear Regression. The evaluated algorithms are systematically assessed using 5-fold cross-validation. In terms of feature importance, SHAP (Shapley Additive exPlanations) values are used to quantify feature importance and characterize associations regarding investor demand, offer price, and listing board. Formal hypothesis testing is conducted via bootstrapped hierarchical regression. Findings SVR-Linear offers better predictive performance (test R 2 = 0.512, CV R 2 = 0.486 ± 0.087), whereas ensemble techniques are prone to overfitting. SHAP analysis shows that the over-subscription ratio (OSR) contributes 40.7%, followed by a knowledge gap (15.6%). Hierarchical regression is used to establish strong moderating roles: investor demand is a significant moderator of the knowledge gap-underpricing relationship (β = 0.0284, p < 0.01), and intense demand (OSR > 30x) dilutes the effect of information asymmetry by 73%. There is marginal moderation in the offer price (β = 0.0198, p = 0.073) and no significant effect by listing board differences (β = 0.0089, p = 0.496). The threshold analysis shows that when investor demand is high, cascading demand effects lead to fundamental valuation concerns and patterns consistent with the predictions of Welch’s (1992) bandwagon theory. Practical implications The findings from this research can benefit a variety of stakeholders. The stockholders can benefit from this research and minimize underpricing. Furthermore, regulatory bodies will consider book-building processes for large IPOs when there is a signal of demand. Additionally, the institutional investors should avoid oversubscription. Furthermore, it can be concluded that precise targeting policies can be achieved by measuring interaction effects. Originality/value The study contributes in three ways, including (1) the first application of SHAP-based interpretability to IPO underpricing, which allows the identification of non-linear moderation effects that are not observed by conventional regression, (2) a rigorous overfitting diagnostic can be evaluated using hierarchical regression with bootstrapped confidence intervals, a feature that identifies the gap between black-box prediction and inferential rigor; and (3) formal statistical validation of SHAP results through bootstrapped hierarchical regression, a property that has not been previously available in financial machine learning studies. Explainable AI, coupled with causal inference, offers a transparent and replicable analytical framework for financial research that requires predictive validity and theoretical elucidation.

Discover Artificial IntelligenceVol. 6(1)
Openalex Percentile: Top 4%
Financial Distress and Bankruptcy Prediction
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.