Training Power Grid Standard LLMs via Token-Adaptive Continual Pretraining and Orthogonal Subspace On-Policy Self-Distillation
Power grid large language models (LLMs) require the efficient acquisition of dense regulatory knowledge while preserving lightweight deployment. We propose a three-stage framework for the efficient training of power grid LLMs via token-adaptive continual pretraining and orthogonal subspace on-policy self-distillation. First, Zone of Proximal Development (ZPD)-guided curriculum selection ranks domain texts using prerequisite knowledge, uncertainty, paraphrase stability, and explanation-induced learning gains, prioritizing samples that are both learnable and informative. Second, Privileged-View Self-Distillation Continual Pretraining (PVD-CPT) estimates token-level knowledge value from asymmetric student–teacher views, strengthens critical numerical, conditional, and normative tokens, and anneals auxiliary weighting and distillation during full-data coverage training. Third, Dual-Teacher Subspace-Guided On-Policy Self-Distillation (DSG-OPSD) evaluates weak and strong views of a frozen teacher along fresh constrained greedy trajectories, estimates a periodically refreshed foundational-correction subspace in the centered full-vocabulary logit space, and retains orthogonal rationale-and-error guidance through spectral-energy and output-uncertainty gating. On a 1024-question single-answer multiple-choice power-grid standards benchmark, where the overall score is question-level accuracy, PVD-CPT improves the score from 19.63 to 32.71. The same DSG-OPSD-trained checkpoint scores 43.26 under direct-answer inference and 49.41 under reasoning inference; the additional 6.15 points are therefore attributable to the inference protocol. On this target-standard benchmark, 49.41 is higher than 44.63 for the strongest evaluated external baseline, Qwen3.5-397B-A17B. Because the external models were evaluated without task-specific adaptation to the 50-standard source collection, this comparison positions target-domain specialization under asymmetric training exposure rather than establishing general-purpose or unseen-standard superiority. The results demonstrate effective domain knowledge acquisition and answer-policy refinement without additional inference-time teachers or auxiliary inputs.
Authors
- Zhiyuan Hu (ORCID: https://orcid.org/0009-0001-5552-7833)
- Chengwei Peng (ORCID: https://orcid.org/0009-0005-8953-2381)
- Zhichao Zhang (ORCID: https://orcid.org/0000-0001-5001-0838)
- Liang Shouyu
- Zhichen Peng
- Zhao Sun
- Jianming Hu
- Zhuojun Cai
Institutions
- China Southern Power Grid (China) (CN)
- Tsinghua University (CN)
Publication Details
- Journal
- Mathematics
- Published
- 2026-09-24
- DOI
- https://doi.org/10.3390/math14193479
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00