Training Power Grid Standard LLMs via Token-Adaptive Continual Pretraining and Orthogonal Subspace On-Policy Self-Distillation

Power grid large language models (LLMs) require the efficient acquisition of dense regulatory knowledge while preserving lightweight deployment. We propose a three-stage framework for the efficient training of power grid LLMs via token-adaptive continual pretraining and orthogonal subspace on-policy self-distillation. First, Zone of Proximal Development (ZPD)-guided curriculum selection ranks domain texts using prerequisite knowledge, uncertainty, paraphrase stability, and explanation-induced learning gains, prioritizing samples that are both learnable and informative. Second, Privileged-View Self-Distillation Continual Pretraining (PVD-CPT) estimates token-level knowledge value from asymmetric student–teacher views, strengthens critical numerical, conditional, and normative tokens, and anneals auxiliary weighting and distillation during full-data coverage training. Third, Dual-Teacher Subspace-Guided On-Policy Self-Distillation (DSG-OPSD) evaluates weak and strong views of a frozen teacher along fresh constrained greedy trajectories, estimates a periodically refreshed foundational-correction subspace in the centered full-vocabulary logit space, and retains orthogonal rationale-and-error guidance through spectral-energy and output-uncertainty gating. On a 1024-question single-answer multiple-choice power-grid standards benchmark, where the overall score is question-level accuracy, PVD-CPT improves the score from 19.63 to 32.71. The same DSG-OPSD-trained checkpoint scores 43.26 under direct-answer inference and 49.41 under reasoning inference; the additional 6.15 points are therefore attributable to the inference protocol. On this target-standard benchmark, 49.41 is higher than 44.63 for the strongest evaluated external baseline, Qwen3.5-397B-A17B. Because the external models were evaluated without task-specific adaptation to the 50-standard source collection, this comparison positions target-domain specialization under asymmetric training exposure rather than establishing general-purpose or unseen-standard superiority. The results demonstrate effective domain knowledge acquisition and answer-policy refinement without additional inference-time teachers or auxiliary inputs.

Authors

Institutions

Publication Details

Journal
Mathematics
Published
2026-09-24
DOI
https://doi.org/10.3390/math14193479
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Training Power Grid Standard LLMs via Token-Adaptive Continual Pretraining and Orthogonal Subspace On-Policy Self-Distillation

Zhiyuan Hu, Chengwei Peng, Zhichao Zhang, Liang Shouyu et al.
Mathematics
Multimodal Machine Learning Applications
article

Training Power Grid Standard LLMs via Token-Adaptive Continual Pretraining and Orthogonal Subspace On-Policy Self-Distillation

Zhiyuan Hu, Chengwei Peng, Zhichao Zhang, Liang Shouyu, Zhichen Peng, Zhao Sun, Jianming Hu, Zhuojun Cai
article en

Abstract

Power grid large language models (LLMs) require the efficient acquisition of dense regulatory knowledge while preserving lightweight deployment. We propose a three-stage framework for the efficient training of power grid LLMs via token-adaptive continual pretraining and orthogonal subspace on-policy self-distillation. First, Zone of Proximal Development (ZPD)-guided curriculum selection ranks domain texts using prerequisite knowledge, uncertainty, paraphrase stability, and explanation-induced learning gains, prioritizing samples that are both learnable and informative. Second, Privileged-View Self-Distillation Continual Pretraining (PVD-CPT) estimates token-level knowledge value from asymmetric student–teacher views, strengthens critical numerical, conditional, and normative tokens, and anneals auxiliary weighting and distillation during full-data coverage training. Third, Dual-Teacher Subspace-Guided On-Policy Self-Distillation (DSG-OPSD) evaluates weak and strong views of a frozen teacher along fresh constrained greedy trajectories, estimates a periodically refreshed foundational-correction subspace in the centered full-vocabulary logit space, and retains orthogonal rationale-and-error guidance through spectral-energy and output-uncertainty gating. On a 1024-question single-answer multiple-choice power-grid standards benchmark, where the overall score is question-level accuracy, PVD-CPT improves the score from 19.63 to 32.71. The same DSG-OPSD-trained checkpoint scores 43.26 under direct-answer inference and 49.41 under reasoning inference; the additional 6.15 points are therefore attributable to the inference protocol. On this target-standard benchmark, 49.41 is higher than 44.63 for the strongest evaluated external baseline, Qwen3.5-397B-A17B. Because the external models were evaluated without task-specific adaptation to the 50-standard source collection, this comparison positions target-domain specialization under asymmetric training exposure rather than establishing general-purpose or unseen-standard superiority. The results demonstrate effective domain knowledge acquisition and answer-policy refinement without additional inference-time teachers or auxiliary inputs.

MathematicsVol. 14(19)
China Southern Power Grid (China) (CN), Tsinghua University (CN)
Quality Education
Openalex Percentile: Top 14%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.