Diagnosing Popularity Collapse in Building Retrofit Shortlists from New York City Energy Audits
This methodological diagnostic tests whether broad building descriptors support retrofit-class shortlists beyond popularity. We mapped 2623 New York City Local Law 87 audit rows to six classes with five decision states, trained on 2019–2022, validated on 2023, and tested on unseen 2024 property groups. Under natural recorded-label visibility, state-aware and equal-information global policies both achieved an observed-positive Recall@3 of 0.832. The state model returned the same Top-3 set for 99.7% of test rows; all evaluated natural-condition neural models missed every recorded insulation and window-upgrade positive. Light-emitting diode (LED) lighting and heating, ventilation, and air conditioning (HVAC) controls comprised 277 of 394 test positives, explaining why high recall did not demonstrate personalization. In a secondary synthetic 10% visibility stress test, excluding unlabelled entries from negative supervision improved recall over naive binary cross-entropy by 0.279; the natural state–naive difference was 0.002 (95% confidence interval (CI) [−0.007, 0.012]). Post hoc capacity, ontology, status-precedence, and stopping sensitivities did not establish reliable superiority over popularity. The benchmark diagnoses label misspecification and nearly constant shortlists.
Authors
- Yunan Zhang (ORCID: https://orcid.org/0009-0006-4567-2849)
- Yanxiao Liu (ORCID: https://orcid.org/0009-0006-1006-7526)
- Shengxi Cao
- Jingjing Fan (ORCID: https://orcid.org/0009-0003-4125-1375)
Institutions
- Zhangjiakou Academy of Agricultural Sciences (CN)
- Computer Emergency Response Team (FR)
- Hebei University of Architecture (CN)
Publication Details
- Journal
- Buildings
- Published
- 2026-09-15
- DOI
- https://doi.org/10.3390/buildings16183666
- Primary Topic
- Sustainable Building Design and Assessment
- Type
- article
- Field-Weighted Citation Impact
- 0.00