Learning instance-level semantic Gaussian representations for occlusion-aware urban scene reconstruction from UAV imagery
UAV-based earth observation provides an efficient and flexible means for detailed urban geospatial modeling and digital twin construction. However, in complex urban environments, foreground occlusions caused by trees, vegetation, and other clutter frequently lead to incomplete structures, contaminated textures, and inconsistent observations across views, thereby reducing the quality and usability of reconstructed urban products. Existing methods still struggle in such scenarios, not only because of severe occlusion itself, but also because they lack a representation at the instance level that can explicitly distinguish target urban structures from occluding interference in 3D space. Without such a representation, Gaussian primitives tend to become entangled across semantic regions, which limits reliable separation between target urban structures and occluders, accurate boundary delineation, and consistent urban scene reconstruction across views from UAV imagery. To address this issue, this paper proposes a semantic Gaussian learning framework at the instance level, enhanced by foundation priors, for urban scene reconstruction from UAV imagery. Specifically, high-level semantic, segmentation, and structural priors extracted from vision foundation models are introduced to guide semantic Gaussian representation learning at the instance level in 3D space. On this basis, an Instance-level Semantic Gaussian Module is designed to inject prior-enhanced semantics into Gaussian primitives, enabling more discriminative representation learning at the instance level. Furthermore, a Prior-guided Interaction Module is developed to establish effective collaboration between visual prior features and 3D Gaussian representations, so that RGB appearance semantics and 3D spatial structure can be jointly encoded within a unified semantic and structural representation. Benefiting from this design, the learned Gaussian representation provides a reliable basis for instance separation, missing-region localization, and consistent urban scene reconstruction across views under severe occlusions. Experimental results on multiple datasets demonstrate that the proposed method consistently outperforms representative baselines in instance separation accuracy, boundary quality, and rendering consistency. The results further show that the proposed representation effectively alleviates instance boundary overlap and category confusion, while improving the completeness, consistency, and usability of urban geospatial products derived from UAV imagery.
Authors
- Zhenfei Ling (ORCID: https://orcid.org/0009-0007-6068-5581)
- Fugui Liu (ORCID: https://orcid.org/0009-0006-9270-7047)
- Ruisheng Wang
- Renzhong Guo
- Tszming Lu
- Kun Zhou
Institutions
- Shenzhen University (CN)
- Urban Planning & Design Institute of Shenzhen (China) (CN)
Publication Details
- Journal
- International Journal of Applied Earth Observation and Geoinformation
- Published
- 2026-09-18
- DOI
- https://doi.org/10.1016/j.jag.2026.105574
- Primary Topic
- Robotics and Sensor-Based Localization
- Type
- article
- Field-Weighted Citation Impact
- 0.00