Neural and Econometric Forecasting of Market Risk: A Comparative Value-at-Risk and Expected-Shortfall Analysis Across Global Equity Markets
This study evaluates whether a deep-learning volatility model improves market-risk measurement relative to established econometric benchmarks. Using daily returns for twelve developed and emerging equity indices from January 2000 to September 2026—with the KSE-100 and IMOEX series taken from the Pakistan Stock Exchange and the Moscow Exchange respectively, because the public feeds for both are truncated—we estimate one-day value-at-risk (VaR) and conditional value-at-risk (CVaR) at the 95% and 99% levels using three approaches: unconditional historical simulation, a GARCH(1,1) model with Student-t innovations and filtered historical simulation (FHS), and a long short-term memory (LSTM) neural network. VaR and CVaR are reported as signed returns, with a negative number denoting a loss; the United States 99% VaR is −3.5% and the corresponding CVaR is −5.0%. Model adequacy is assessed with the Kupiec unconditional-coverage and Christoffersen conditional-coverage backtests—decomposed into their independence (LR_ind) and coverage (LR_uc) components—and volatility-forecast accuracy is compared out-of-sample using the QLIKE robust loss. Four findings emerge. First, static historical VaR attains correct average coverage yet fails the conditional-coverage test in eleven of twelve markets; the failure is driven almost entirely by the independence component, confirming that its exceedances cluster in high-volatility episodes rather than reflecting incorrect average coverage. Second, the GARCH-FHS model, which conditions the risk estimate on time-varying volatility, restores conditional coverage in all twelve markets, including on a backtest window matched exactly to the LSTM’s held-out test period, where it passes in eight of twelve markets against two of twelve for the LSTM. Third, the LSTM is correctly signed but poorly calibrated: with the plug-in FHS construction it jointly passes both backtests in only two of twelve markets, with breach rates well below the nominal 5% in most of the rest, a pattern consistent with a significant drift between the training-window and test-window standardised-residual distributions (Kolmogorov–Smirnov p < 0.05 in eight of twelve markets). Roughly half of that over-conservatism is attributable to the two-step plug-in design rather than to the network: retraining the same architecture directly on the tail quantiles with a composite pinball loss raises the mean breach rate from 2.8% to 4.3% against a nominal 5% but still yields joint coverage in only six of twelve markets, against eight of twelve for a window-matched GARCH-FHS. Fourth, a Diebold–Mariano test on the QLIKE loss finds GARCH significantly more accurate than the LSTM in ten of twelve markets at the 5% level and in nine of twelve after a Holm–Bonferroni correction for multiple testing; the LSTM is not significantly more accurate than GARCH in any market, before or after correction—nor in any of the 320 additional training runs reported below. These results are reported for a specific neural implementation, and we are deliberate about the scope of the claim they support: we do not find evidence that a single-layer LSTM trained with a Huber loss on the next-day absolute return and converted to VaR through a plug-in FHS residual distribution is competitive with GARCH(1,1)-t FHS in any of the twelve markets we study, under a randomised search over six training hyperparameters comprising 40 configurations in each of five markets and across ten random seeds per market in all twelve. The ranking is unchanged when the LSTM is instead trained directly on the quantile it is judged on with a composite pinball loss, when the Gaussian absolute-return conversion is replaced by its empirical counterpart, and when a Parkinson high–low range proxy replaces the squared return in the loss function. We do not claim, and our design cannot establish, that no neural architecture can beat GARCH for this task. The decisive gain in VaR adequacy comes from conditioning on dynamic volatility, and within the conditional-model class a parsimonious GARCH(1,1) is a demanding benchmark. We discuss implications for regulatory backtesting, tail-risk management, and the design of hybrid econometric-neural systems and report the full replication pipeline (data-download, modelling, and backtesting code) alongside this submission.
Authors
- Zeeshan Ahmed (ORCID: https://orcid.org/0000-0001-8388-3168)
- HASSAN Arshad
- Raima Amjad (ORCID: https://orcid.org/0009-0002-0748-3888)
Institutions
- Capital University of Science and Technology (PK)
Publication Details
- Journal
- Risks
- Published
- 2026-09-30
- DOI
- https://doi.org/10.3390/risks14100228
- Primary Topic
- Financial Risk and Volatility Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00