A data-efficient and physics-guided agent-based deep reinforcement learning framework for cost-aware multi-layer shielding design in compact nuclear reactors
Compact nuclear energy systems such as small modular reactors and microreactors are emerging as versatile, carbon-free, and resilient energy providers. To enable these reactors to be economically viable, it is essential to design compact and cost-effective radiation shielding solutions that go beyond traditional shielding methods. This study tackles the shielding challenge by proposing a data-efficient, physics-guided, agent-based, and cost-aware deep reinforcement learning framework (D-PAC) for multi-layer shielding design, built on a Rainbow deep Q-network backbone. In the simplified Savannah primary-shield model, D-PAC identifies a configuration that reduces shield volume by approximately 13%, total mass by approximately 9%, and estimated material cost by approximately 11%, while maintaining the calculated external dose below the prescribed limit. The related source and target shielding benchmark studies further evaluate information reuse across tasks. In the target shielding case, teacher-guided configurations reach matched reward thresholds earlier than the random-initialization control. The distilled policy reaches the matched genetic algorithm reward target with more than a tenfold reduction in recorded evaluations relative to the direct genetic algorithm trace. Together, these studies demonstrate an integrated workflow for combining prior shielding knowledge, learning-based search, and high-fidelity radiation-transport evaluation.
Authors
- Kai Tan
- Fan Zhang (ORCID: https://orcid.org/0000-0002-4974-3329)
Institutions
- Georgia Institute of Technology (US)
Publication Details
- Journal
- Engineering Applications of Artificial Intelligence
- Published
- 2026-09-24
- DOI
- https://doi.org/10.1016/j.engappai.2026.116306
- Primary Topic
- Nuclear reactor physics and engineering
- Type
- article
- Field-Weighted Citation Impact
- 0.00