A Comparative Analysis of Four Semantic Segmentation Models for Classification of Building Structural Elements in Highly Occluded Environments

Semantic segmentation of point clouds plays a critical role in automatic Scan-to-BIM workflows. This study evaluates the performance of two Multi-Layer-Perceptron (MLP)-based deep learning networks (PointNet++, PointNeXt-XL) and two transformer-based networks (Point Transformer V1 (PTv 1), Point Transformer V3 (PTv 3)) for the identification of building structural elements in highly occluded environments. The models are trained and tested on three LiDAR datasets of reinforced concrete structures including office buildings and a multi-storey carpark. Results show that transformer-based networks significantly outperform MLP-based architectures under heavy occlusions. Within the transformer-based models, PTv 1 achieved the highest Overall Accuracy (OA) at 92.62% while PTv 3 provided the best balance across all classes, particularly for beam and clutter classes due to its larger receptive field and enhanced geometric encoding. Because segmentation approaches can compensate for errors in ceiling, floor, and column class identification, we recommend PTv 3 for Scan-to-BIM applications focused on structural modelling.

Authors

Institutions

Publication Details

Journal
ISPRS annals of the photogrammetry, remote sensing and spatial information sciences
Published
2026-09-28
DOI
https://doi.org/10.5194/isprs-annals-xii-4-w1-2026-1-2026
Primary Topic
3D Surveying and Cultural Heritage
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

A Comparative Analysis of Four Semantic Segmentation Models for Classification of Building Structural Elements in Highly Occluded Environments

Davood Shojaei, Martin Tomko, Mojtaba Akhoundi Khezrabad
ISPRS annals of the photogrammetry, remote sensing and spatial information sciences
3D Surveying and Cultural Heritage
article

A Comparative Analysis of Four Semantic Segmentation Models for Classification of Building Structural Elements in Highly Occluded Environments

Davood Shojaei, Martin Tomko, Mojtaba Akhoundi Khezrabad
article en

Abstract

Semantic segmentation of point clouds plays a critical role in automatic Scan-to-BIM workflows. This study evaluates the performance of two Multi-Layer-Perceptron (MLP)-based deep learning networks (PointNet++, PointNeXt-XL) and two transformer-based networks (Point Transformer V1 (PTv 1), Point Transformer V3 (PTv 3)) for the identification of building structural elements in highly occluded environments. The models are trained and tested on three LiDAR datasets of reinforced concrete structures including office buildings and a multi-storey carpark. Results show that transformer-based networks significantly outperform MLP-based architectures under heavy occlusions. Within the transformer-based models, PTv 1 achieved the highest Overall Accuracy (OA) at 92.62% while PTv 3 provided the best balance across all classes, particularly for beam and clutter classes due to its larger receptive field and enhanced geometric encoding. Because segmentation approaches can compensate for errors in ceiling, floor, and column class identification, we recommend PTv 3 for Scan-to-BIM applications focused on structural modelling.

ISPRS annals of the photogrammetry, remote sensing and spatial information sciencesVol. XII-4/W1-2026(0)
The University of Melbourne (AU)
Sustainable cities and communities
Openalex Percentile: Top 9%
3D Surveying and Cultural Heritage
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.