Evaluating Grounded Alignment of LLM Explanations in Financial Risk Early Warning: A Prediction–Explanation Separation Framework
Purpose: This study proposes a large language model (LLM) framework that separates prediction from explanation for early warning of high-volatility regimes in the Korean stock market, in which a dedicated model issues the forecast and an LLM produces auditable, evidence-grounded explanations. The goal is to evaluate and manage the grounded alignment of LLM explanations rather than to use the LLM as a forecaster.Methods: Daily Korea Composite Stock Price Index (KOSPI) and global-market data from January 2015 to May 2026 were used to construct 35 derived features; global variables were aligned at a one-trading-day lag to remove look-ahead bias, and a five-day embargo was applied at split boundaries. Under a strict chronological split, seven models were compared and the main predictor was selected by the area under the precision– recall curve (PR-AUC) on the validation set. The predictor estimates whether next-five-day realized volatility exceeds the top-25% training threshold; its predicted probability, SHapley Additive exPlanations (SHAP) attributions, a market-state summary, and a risk-factor taxonomy are assembled into a Structured Evidence Packet (SEP) that is the LLM's only input. Explanation quality was assessed by a Grounded Evidence Alignment Score (GEAS) and repetition stability, with input ablation (P1/P3/P4) and deletion/injection interventions.Results: XGBoost was selected as the main predictor (validation PR-AUC 0.760), with a test area under the receiver operating characteristic curve (ROC-AUC) of 0.665 and a PR-AUC of 0.782; pairwise model differences were not significant under a moving-block bootstrap. GEAS rose from 0.183 (no SHAP) to 1.000/0.998 when SHAP evidence was provided (paired Wilcoxon, Holm-adjusted p<1e-21), with no significant P3–P4 difference (a ceiling effect). On a common set of 120 sampled test dates, the dedicated model attained PR-AUC 0.799 versus 0.615 for direct LLM prediction (a reference comparison, because the two predictors received inputs of different information scope). Removing the top SHAP cue eliminated its mention (0%), whereas injecting an irrelevant distractor as the top cue produced 100% adoption.Conclusion: GEAS captures adherence to the provided evidence boundary rather than the LLM's internal reasoning faithfulness, and the LLM does not independently validate evidence quality. The findings support using LLMs as grounded, auditable explanation layers rather than autonomous forecasters, with system reliability contingent on the quality of the predictor's evidence.
Authors
- Nak Hyun Jung (ORCID: https://orcid.org/0009-0007-7390-1714)
- SakJae Lim
Institutions
- Seoul School of Integrated Sciences and Technologies (KR)
Publication Details
- Journal
- Journal of the Korean society for quality management
- Published
- 2026-09-29
- DOI
- https://doi.org/10.7469/jksqm.2026.54.3.457
- Primary Topic
- Stock Market Forecasting Methods
- Type
- article
- Field-Weighted Citation Impact
- 0.00