Explainable Reinforcement Learning for Transmission-Corridor Reinforcement under Renewable Integration: An Austrian Case Study
Transmission expansion planning must balance investment cost, network loading, and renewable-energy integration under uncertain operating conditions. This study investigates whether explainable reinforcement learning can support the allocation of additional capacity to existing transmission corridors in a stylised Austrian case. An Austrian network derived from PyPSA-Eur was combined with hourly realised-load and renewable-availability profiles for 2015--2024 and with regional wind, photovoltaic, and peak-load assumptions representing 2030 and 2040. A supplementary 2035 input was obtained by linear interpolation. The sequential task comprised four reinforcement decisions within a 24-hour episode, a 500~MW aggregate envelope, and 60 screened candidate corridors. Scalar proximal policy optimisation (PPO) and preference-conditioned multi-objective PPO (MO-PPO) were trained with five independent seeds in a computationally efficient linearised proxy. Policies were evaluated on ten shared held-out windows, against deterministic strategies and a multi-snapshot linear reinforcement reference, and in a stricter PyPSA dispatch model. Policy explanations combined output-specific permutation Shapley attribution, grouped-holdout ridge surrogates, and a capacity-preserving corridor intervention. In the proxy, PPO reduced grid stress by 22.0% and renewable curtailment by 12.1% relative to zero reinforcement, but targeted deterministic strategies achieved larger improvements with less capacity. In strict evaluation, reinforcement investment charges exceeded the operating savings of both learned methods, and PPO and MO-PPO were not statistically distinguishable. At the central cost assumption, the linear reference selected 305.58MW on three Salzburg-area corridors for the 2040 input and no reinforcement for the 2030 or interpolated 2035 inputs. In contrast, the transferred learned policies proposed nearly invariant totals of approximately 382-385MW. Shapley estimates were repeatable within a trained policy but varied substantially across seeds; surrogate fidelity depended on the policy output. Reallocating capacity away from the three reference-selected corridors increased operating cost and curtailment in point estimates, although the total-cost interval included zero. The study demonstrates an auditable framework for detecting proxy-to-strict, cross-scenario, statistical, and explanation-stability gaps. It does not establish superiority of reinforcement learning or justify a concrete Austrian transmission project.
Authors
- Anna Till
- Rosana De Oliveira Gomes
Institutions
- University of Applied Sciences Technikum Wien (AT)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-14
- DOI
- https://doi.org/10.5281/zenodo.22755273
- Primary Topic
- Electric Power System Optimization
- Type
- preprint