Multi-Objective Reinforcement Learning for Economic Optimization of Photovoltaic Cleaning and Cooling Schedules

Photovoltaic (PV) output in arid regions is reduced by dust soiling and elevated module temperatures, while cleaning and spray cooling consume water, energy, and labor. Scheduling these interventions is therefore an economic problem that fixed intervals and simple thresholds address imperfectly. We develop a reproducible simulator for five Moroccan sites across four climate regimes, using Copernicus Atmosphere Monitoring Service (CAMS) dust aerosol and ERA5-Land meteorological reanalysis for 2004–2025. The model’s single soiling parameter is calibrated against published field rates; the modeled rates for all five sites fall within their target ranges. Cleaning and cooling costs are calculated in Moroccan dirhams from their physical components. We formulate maintenance as a Markov decision process and train separate Proximal Policy Optimization policies across a range of weights assigned to energy and cost. An exhaustive sweep of 21 scripted policies per site reveals a flat response surface: the best policy in the evaluated grid cleans when the soiling fraction exceeds approximately 0.10 and disables spray cooling at every site. Learned policies capture 52–69% of the grid-best improvement over no action, compared with 62.2–96.6% for the two interpretable scripted rules. The shortfall persists across three reinforcement learning algorithms; probing the networks identifies the placement of the learned cleaning trigger as a key mechanism. In a separate 30-city analysis, the value of scripted maintenance rises with aridity, from 6.2% to 50.9%. Absolute monetary benefits depend on assumed constants, particularly the soiling ceiling, which accounts for an 89.6% sensitivity span. The study provides a calibrated comparison of scheduling policies and a diagnosis of why the learned policies underperform the scripted rules under the modeled conditions.

Authors

Institutions

Publication Details

Journal
Solar
Published
2026-10-06
DOI
https://doi.org/10.3390/solar6050070
Primary Topic
Solar Thermal and Photovoltaic Systems
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Multi-Objective Reinforcement Learning for Economic Optimization of Photovoltaic Cleaning and Cooling Schedules

Mohamed-Amine Babay, Youssef Ait El Kadi, Ayoub Boussiri
Solar
Solar Thermal and Photovoltaic Systems
article

Multi-Objective Reinforcement Learning for Economic Optimization of Photovoltaic Cleaning and Cooling Schedules

Mohamed-Amine Babay, Youssef Ait El Kadi, Ayoub Boussiri
article en

Abstract

Photovoltaic (PV) output in arid regions is reduced by dust soiling and elevated module temperatures, while cleaning and spray cooling consume water, energy, and labor. Scheduling these interventions is therefore an economic problem that fixed intervals and simple thresholds address imperfectly. We develop a reproducible simulator for five Moroccan sites across four climate regimes, using Copernicus Atmosphere Monitoring Service (CAMS) dust aerosol and ERA5-Land meteorological reanalysis for 2004–2025. The model’s single soiling parameter is calibrated against published field rates; the modeled rates for all five sites fall within their target ranges. Cleaning and cooling costs are calculated in Moroccan dirhams from their physical components. We formulate maintenance as a Markov decision process and train separate Proximal Policy Optimization policies across a range of weights assigned to energy and cost. An exhaustive sweep of 21 scripted policies per site reveals a flat response surface: the best policy in the evaluated grid cleans when the soiling fraction exceeds approximately 0.10 and disables spray cooling at every site. Learned policies capture 52–69% of the grid-best improvement over no action, compared with 62.2–96.6% for the two interpretable scripted rules. The shortfall persists across three reinforcement learning algorithms; probing the networks identifies the placement of the learned cleaning trigger as a key mechanism. In a separate 30-city analysis, the value of scripted maintenance rises with aridity, from 6.2% to 50.9%. Absolute monetary benefits depend on assumed constants, particularly the soiling ceiling, which accounts for an 89.6% sensitivity span. The study provides a calibrated comparison of scheduling policies and a diagnosis of why the learned policies underperform the scripted rules under the modeled conditions.

SolarVol. 6(5)
Université Ibn Zohr (MA), Université Sultan Moulay Slimane (MA)
Openalex Percentile: Top 33%
Solar Thermal and Photovoltaic Systems
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.