Hybrid constraint programming and deep reinforcement learning for automated retail facility layout planning
Convenience-store layout planning is a knowledge-intensive design task that involves heterogeneous facility objects and complex spatial constraints; however, it is still largely performed by human designers using experience-driven heuristics, resulting in inconsistent quality, limited scalability, and substantial time and labor costs. This article proposes a hybrid framework for automated convenience-store layout planning that integrates Deep Reinforcement Learning (DRL) and Constraint Programming (CP). The problem is formulated as a sequential decision-making process in which a DRL agent determines the placement order of facility objects, and a CP module searches for feasible locations subject to spatial and operational constraints. The framework generates candidate store configurations within seconds and is intended to support rapid alternative exploration; its effect on human design effort was not measured in this study. Experiments on 100 validation instances show that the proposed method improves the Exponentially Weighted Scoring Metric (EWSM) by 1.60% over random multi-start search. Compared with SequenceGA, the observed mean EWSM values are similar (1.0187 versus 1.0173), while the proposed method requires approximately half the mean runtime; the paired sign test does not establish a significant difference in EWSM. In a controlled ablation that changes only learned versus uniform class selection, the learned policy improves mean EWSM by 0.63% (95% bootstrap confidence interval: 0.06%–1.37%) and substantially improves lower-tail robustness, although rank- and sign-based paired tests do not establish uniform per-instance superiority. In the IntegratedCP-Small benchmark, which uses a finite-candidate integrated CP-SAT formulation on 40 reduced instances, CP–DRL matches the optimal surplus objective of the prespecified 30-unit candidate pool in 35 cases, with a mean relative gap of 2.17% (case-clustered 95% bootstrap confidence interval: 0.00%–4.83%); a 15-unit grid sensitivity yields 34 matches. These results position CP–DRL as a learning-guided constructive heuristic rather than a globally optimal solver. A structured multimodal evaluation, compared descriptively with ratings from seven human experts, is retained only as an optional post hoc assessment tool for human-centered layout qualities.
Authors
- Chien‐Liang Liu (ORCID: https://orcid.org/0000-0002-2724-7199)
- Yung-Chu Lin
- Ming-Han Chang (ORCID: https://orcid.org/0009-0009-4363-9959)
Institutions
- National Yang Ming Chiao Tung University (TW)
Publication Details
- Journal
- Advanced Engineering Informatics
- Published
- 2026-09-11
- DOI
- https://doi.org/10.1016/j.aei.2026.105201
- Primary Topic
- Advanced Manufacturing and Logistics Optimization
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Science and Technology Council