Reinforcement learning–enhanced population-based multi-objective ore blending with iot-observed grade uncertainty in open-pit mines

Abstract Open-pit ore blending must simultaneously control grade-target compliance, economic value, and operating cost with noisy, partial-grade observations from the Mining 4.0 sensing infrastructure. Building on recent population-based multi-objective reinforcement learning (MORL) architectures for ore blending, this work presents an in-depth empirical study of sequential decision-making under IoT-observed grade uncertainty, involving three competing objectives: grade deviation, net present value (NPV), and cost. To overcome preference collapse and insufficient Pareto coverage in single-policy MORL, we systematically extend the specialist-bank approach by introducing uncertainty-weighted multi-objective gradients, a curriculum noise annealing protocol tied to sensor characteristics, and plant-level operational constraints including crusher throughput and haulage costs. We evaluate this extended architecture against single-policy MORL, rolling-horizon evolutionary control, and static evolutionary baselines on open benchmark mine instances under multiple uncertainty levels. The population MORL achieves a higher peak NPV than static NSGA-II on the primary mine instance, matches the performance of rolling-horizon evolutionary control at three orders of magnitude lower inference latency, and generalizes to an external mine instance (Newman1) without retraining, demonstrating practical transfer to a different orebody distribution. These results suggest that population-based MORL is a viable path to realizing real-time, preference-aware blending control in the face of sensor uncertainty and highlight the importance of empirical, constraint-aware analysis when deploying these architectures alongside offline evolutionary optimization.

Authors

Institutions

Publication Details

Journal
Journal of Engineering and Applied Science
Published
2026-08-25
DOI
https://doi.org/10.1186/s44147-026-01185-2
Primary Topic
Mining Techniques and Economics
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Reinforcement learning–enhanced population-based multi-objective ore blending with iot-observed grade uncertainty in open-pit mines

Azamat Umirzoqov, Aidar Kuttybayev, Ainash Kainazarova, Shokhjakhon Abdufattokhov et al.
Journal of Engineering and Applied Science
Mining Techniques and Economics
article

Reinforcement learning–enhanced population-based multi-objective ore blending with iot-observed grade uncertainty in open-pit mines

Azamat Umirzoqov, Aidar Kuttybayev, Ainash Kainazarova, Shokhjakhon Abdufattokhov, Galymzhan Samenov, Arystan Kozhantov, Shuhratulla Ochilov, Махфуза Тухтаева, Suhbat Norinov, Kazi Bizhanov, Ryumduk Oh
article en

Abstract

Abstract Open-pit ore blending must simultaneously control grade-target compliance, economic value, and operating cost with noisy, partial-grade observations from the Mining 4.0 sensing infrastructure. Building on recent population-based multi-objective reinforcement learning (MORL) architectures for ore blending, this work presents an in-depth empirical study of sequential decision-making under IoT-observed grade uncertainty, involving three competing objectives: grade deviation, net present value (NPV), and cost. To overcome preference collapse and insufficient Pareto coverage in single-policy MORL, we systematically extend the specialist-bank approach by introducing uncertainty-weighted multi-objective gradients, a curriculum noise annealing protocol tied to sensor characteristics, and plant-level operational constraints including crusher throughput and haulage costs. We evaluate this extended architecture against single-policy MORL, rolling-horizon evolutionary control, and static evolutionary baselines on open benchmark mine instances under multiple uncertainty levels. The population MORL achieves a higher peak NPV than static NSGA-II on the primary mine instance, matches the performance of rolling-horizon evolutionary control at three orders of magnitude lower inference latency, and generalizes to an external mine instance (Newman1) without retraining, demonstrating practical transfer to a different orebody distribution. These results suggest that population-based MORL is a viable path to realizing real-time, preference-aware blending control in the face of sensor uncertainty and highlight the importance of empirical, constraint-aware analysis when deploying these architectures alongside offline evolutionary optimization.

Journal of Engineering and Applied ScienceVol. 73(1)
L. N. Gumilyov Eurasian National University (KZ), Korea National University of Transportation (KR), Satbayev University (KZ), Tashkent State Technical University named after Islam Karimov (UZ), D. Serikbayev East Kazakhstan State Technical University (KZ), Turin Polytechnic University (UZ), National University of Uzbekistan (UZ), Westminster International University in Tashkent (UZ), Keihin (Japan) (JP)
Industry, innovation and infrastructure
Openalex Percentile: Top 14%
Mining Techniques and Economics
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.