Evaluating Grounded Alignment of LLM Explanations in Financial Risk Early Warning: A Prediction–Explanation Separation Framework

Purpose: This study proposes a large language model (LLM) framework that separates prediction from explanation for early warning of high-volatility regimes in the Korean stock market, in which a dedicated model issues the forecast and an LLM produces auditable, evidence-grounded explanations. The goal is to evaluate and manage the grounded alignment of LLM explanations rather than to use the LLM as a forecaster.Methods: Daily Korea Composite Stock Price Index (KOSPI) and global-market data from January 2015 to May 2026 were used to construct 35 derived features; global variables were aligned at a one-trading-day lag to remove look-ahead bias, and a five-day embargo was applied at split boundaries. Under a strict chronological split, seven models were compared and the main predictor was selected by the area under the precision– recall curve (PR-AUC) on the validation set. The predictor estimates whether next-five-day realized volatility exceeds the top-25% training threshold; its predicted probability, SHapley Additive exPlanations (SHAP) attributions, a market-state summary, and a risk-factor taxonomy are assembled into a Structured Evidence Packet (SEP) that is the LLM's only input. Explanation quality was assessed by a Grounded Evidence Alignment Score (GEAS) and repetition stability, with input ablation (P1/P3/P4) and deletion/injection interventions.Results: XGBoost was selected as the main predictor (validation PR-AUC 0.760), with a test area under the receiver operating characteristic curve (ROC-AUC) of 0.665 and a PR-AUC of 0.782; pairwise model differences were not significant under a moving-block bootstrap. GEAS rose from 0.183 (no SHAP) to 1.000/0.998 when SHAP evidence was provided (paired Wilcoxon, Holm-adjusted p<1e-21), with no significant P3–P4 difference (a ceiling effect). On a common set of 120 sampled test dates, the dedicated model attained PR-AUC 0.799 versus 0.615 for direct LLM prediction (a reference comparison, because the two predictors received inputs of different information scope). Removing the top SHAP cue eliminated its mention (0%), whereas injecting an irrelevant distractor as the top cue produced 100% adoption.Conclusion: GEAS captures adherence to the provided evidence boundary rather than the LLM's internal reasoning faithfulness, and the LLM does not independently validate evidence quality. The findings support using LLMs as grounded, auditable explanation layers rather than autonomous forecasters, with system reliability contingent on the quality of the predictor's evidence.

Authors

Institutions

Publication Details

Journal
Journal of the Korean society for quality management
Published
2026-09-29
DOI
https://doi.org/10.7469/jksqm.2026.54.3.457
Primary Topic
Stock Market Forecasting Methods
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evaluating Grounded Alignment of LLM Explanations in Financial Risk Early Warning: A Prediction–Explanation Separation Framework

Nak Hyun Jung, SakJae Lim
Journal of the Korean society for quality management
Stock Market Forecasting Methods
article

Evaluating Grounded Alignment of LLM Explanations in Financial Risk Early Warning: A Prediction–Explanation Separation Framework

Nak Hyun Jung, SakJae Lim
article en

Abstract

Purpose: This study proposes a large language model (LLM) framework that separates prediction from explanation for early warning of high-volatility regimes in the Korean stock market, in which a dedicated model issues the forecast and an LLM produces auditable, evidence-grounded explanations. The goal is to evaluate and manage the grounded alignment of LLM explanations rather than to use the LLM as a forecaster.Methods: Daily Korea Composite Stock Price Index (KOSPI) and global-market data from January 2015 to May 2026 were used to construct 35 derived features; global variables were aligned at a one-trading-day lag to remove look-ahead bias, and a five-day embargo was applied at split boundaries. Under a strict chronological split, seven models were compared and the main predictor was selected by the area under the precision– recall curve (PR-AUC) on the validation set. The predictor estimates whether next-five-day realized volatility exceeds the top-25% training threshold; its predicted probability, SHapley Additive exPlanations (SHAP) attributions, a market-state summary, and a risk-factor taxonomy are assembled into a Structured Evidence Packet (SEP) that is the LLM's only input. Explanation quality was assessed by a Grounded Evidence Alignment Score (GEAS) and repetition stability, with input ablation (P1/P3/P4) and deletion/injection interventions.Results: XGBoost was selected as the main predictor (validation PR-AUC 0.760), with a test area under the receiver operating characteristic curve (ROC-AUC) of 0.665 and a PR-AUC of 0.782; pairwise model differences were not significant under a moving-block bootstrap. GEAS rose from 0.183 (no SHAP) to 1.000/0.998 when SHAP evidence was provided (paired Wilcoxon, Holm-adjusted p<1e-21), with no significant P3–P4 difference (a ceiling effect). On a common set of 120 sampled test dates, the dedicated model attained PR-AUC 0.799 versus 0.615 for direct LLM prediction (a reference comparison, because the two predictors received inputs of different information scope). Removing the top SHAP cue eliminated its mention (0%), whereas injecting an irrelevant distractor as the top cue produced 100% adoption.Conclusion: GEAS captures adherence to the provided evidence boundary rather than the LLM's internal reasoning faithfulness, and the LLM does not independently validate evidence quality. The findings support using LLMs as grounded, auditable explanation layers rather than autonomous forecasters, with system reliability contingent on the quality of the predictor's evidence.

Journal of the Korean society for quality managementVol. 54(3)
Seoul School of Integrated Sciences and Technologies (KR)
Openalex Percentile: Top 7%
Stock Market Forecasting Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.