Explainable Reinforcement Learning for Transmission-Corridor Reinforcement under Renewable Integration: An Austrian Case Study

Transmission expansion planning must balance investment cost, network loading, and renewable-energy integration under uncertain operating conditions. This study investigates whether explainable reinforcement learning can support the allocation of additional capacity to existing transmission corridors in a stylised Austrian case. An Austrian network derived from PyPSA-Eur was combined with hourly realised-load and renewable-availability profiles for 2015--2024 and with regional wind, photovoltaic, and peak-load assumptions representing 2030 and 2040. A supplementary 2035 input was obtained by linear interpolation. The sequential task comprised four reinforcement decisions within a 24-hour episode, a 500~MW aggregate envelope, and 60 screened candidate corridors. Scalar proximal policy optimisation (PPO) and preference-conditioned multi-objective PPO (MO-PPO) were trained with five independent seeds in a computationally efficient linearised proxy. Policies were evaluated on ten shared held-out windows, against deterministic strategies and a multi-snapshot linear reinforcement reference, and in a stricter PyPSA dispatch model. Policy explanations combined output-specific permutation Shapley attribution, grouped-holdout ridge surrogates, and a capacity-preserving corridor intervention. In the proxy, PPO reduced grid stress by 22.0% and renewable curtailment by 12.1% relative to zero reinforcement, but targeted deterministic strategies achieved larger improvements with less capacity. In strict evaluation, reinforcement investment charges exceeded the operating savings of both learned methods, and PPO and MO-PPO were not statistically distinguishable. At the central cost assumption, the linear reference selected 305.58MW on three Salzburg-area corridors for the 2040 input and no reinforcement for the 2030 or interpolated 2035 inputs. In contrast, the transferred learned policies proposed nearly invariant totals of approximately 382-385MW. Shapley estimates were repeatable within a trained policy but varied substantially across seeds; surrogate fidelity depended on the policy output. Reallocating capacity away from the three reference-selected corridors increased operating cost and curtailment in point estimates, although the total-cost interval included zero. The study demonstrates an auditable framework for detecting proxy-to-strict, cross-scenario, statistical, and explanation-stability gaps. It does not establish superiority of reinforcement learning or justify a concrete Austrian transmission project.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-14
DOI
https://doi.org/10.5281/zenodo.22755273
Primary Topic
Electric Power System Optimization
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Explainable Reinforcement Learning for Transmission-Corridor Reinforcement under Renewable Integration: An Austrian Case Study

Anna Till, Rosana De Oliveira Gomes
Zenodo (CERN European Organization for Nuclear Research)
Electric Power System Optimization
preprint

Explainable Reinforcement Learning for Transmission-Corridor Reinforcement under Renewable Integration: An Austrian Case Study

Anna Till, Rosana De Oliveira Gomes
preprint en

Abstract

Transmission expansion planning must balance investment cost, network loading, and renewable-energy integration under uncertain operating conditions. This study investigates whether explainable reinforcement learning can support the allocation of additional capacity to existing transmission corridors in a stylised Austrian case. An Austrian network derived from PyPSA-Eur was combined with hourly realised-load and renewable-availability profiles for 2015--2024 and with regional wind, photovoltaic, and peak-load assumptions representing 2030 and 2040. A supplementary 2035 input was obtained by linear interpolation. The sequential task comprised four reinforcement decisions within a 24-hour episode, a 500~MW aggregate envelope, and 60 screened candidate corridors. Scalar proximal policy optimisation (PPO) and preference-conditioned multi-objective PPO (MO-PPO) were trained with five independent seeds in a computationally efficient linearised proxy. Policies were evaluated on ten shared held-out windows, against deterministic strategies and a multi-snapshot linear reinforcement reference, and in a stricter PyPSA dispatch model. Policy explanations combined output-specific permutation Shapley attribution, grouped-holdout ridge surrogates, and a capacity-preserving corridor intervention. In the proxy, PPO reduced grid stress by 22.0% and renewable curtailment by 12.1% relative to zero reinforcement, but targeted deterministic strategies achieved larger improvements with less capacity. In strict evaluation, reinforcement investment charges exceeded the operating savings of both learned methods, and PPO and MO-PPO were not statistically distinguishable. At the central cost assumption, the linear reference selected 305.58MW on three Salzburg-area corridors for the 2040 input and no reinforcement for the 2030 or interpolated 2035 inputs. In contrast, the transferred learned policies proposed nearly invariant totals of approximately 382-385MW. Shapley estimates were repeatable within a trained policy but varied substantially across seeds; surrogate fidelity depended on the policy output. Reallocating capacity away from the three reference-selected corridors increased operating cost and curtailment in point estimates, although the total-cost interval included zero. The study demonstrates an auditable framework for detecting proxy-to-strict, cross-scenario, statistical, and explanation-stability gaps. It does not establish superiority of reinforcement learning or justify a concrete Austrian transmission project.

Zenodo (CERN European Organization for Nuclear Research)
University of Applied Sciences Technikum Wien (AT)
Affordable and clean energy
Electric Power System Optimization
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.