SceneContextNet: an integrated multi-stage deep learning framework for context-aware scene understanding using multi-object detection and spatial relationship modelling

Deep learning, especially CNN- and transformer-based approaches, has significantly impacted computer vision techniques such as object detection, segmentation, and scene classification. Despite the many successful scene understanding pipelines currently in use, most still rely primarily on global image features or object-level predictions and have little capacity to explicitly model object relationships, spatial arrangements, and structured contextual dependencies. This limits their performance in complex scenes with multiple objects in different semantic layouts. To address this limitation, this paper introduces SceneContextNet. This relationship-guided scene understanding framework transforms detected objects into a structured scene graph and reasons over object–relation representations with context. The novel idea of the proposed approach is to employ a spatial–semantic graph-sequence reasoning strategy, which encodes the object features and pairwise relationship descriptors as nodes and edges, respectively, in the graph; the graph is then convolved to refine the features of objects and relations, and spatial sequence of the objects and relations is further modelled by Bi-LSTM; the attention-based object–relation feature fusion module aggregates the features of the objects and relations. SceneContextNet explicitly connects object detection, relationship modelling, graph construction, context reasoning, and interpretable prediction in one unified flow, unlike traditional module-stacking methods. We integrate explainability by visualising influential image regions, objects, and relationships through Grad-CAM, LIME, and graph-edge visualisation. Experimental results on ADE20K, Visual Genome and SUN RGB-D indicate that SceneContextNet can reach an accuracy of 98.62%, an F1-score of 98.51%, and an mAP of 98.30%, which are superior to those of CNN-, transformer-, and graph-based baselines. Further, cross-dataset evaluation shows strong generalisation across different scene layouts and object distributions. The proposed framework can be utilised in autonomous systems, intelligent surveillance, assistive vision, and smart environments where accurate and interpretable scene understanding is needed.

Authors

Institutions

Publication Details

Journal
Discover Computing
Published
2026-09-28
DOI
https://doi.org/10.1007/s10791-026-10615-x
Primary Topic
Multimodal Machine Learning Applications
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

SceneContextNet: an integrated multi-stage deep learning framework for context-aware scene understanding using multi-object detection and spatial relationship modelling

Venubabu Rachapudi, M. Srividya
Discover Computing
Multimodal Machine Learning Applications
article

SceneContextNet: an integrated multi-stage deep learning framework for context-aware scene understanding using multi-object detection and spatial relationship modelling

Venubabu Rachapudi, M. Srividya
article en

Abstract

Deep learning, especially CNN- and transformer-based approaches, has significantly impacted computer vision techniques such as object detection, segmentation, and scene classification. Despite the many successful scene understanding pipelines currently in use, most still rely primarily on global image features or object-level predictions and have little capacity to explicitly model object relationships, spatial arrangements, and structured contextual dependencies. This limits their performance in complex scenes with multiple objects in different semantic layouts. To address this limitation, this paper introduces SceneContextNet. This relationship-guided scene understanding framework transforms detected objects into a structured scene graph and reasons over object–relation representations with context. The novel idea of the proposed approach is to employ a spatial–semantic graph-sequence reasoning strategy, which encodes the object features and pairwise relationship descriptors as nodes and edges, respectively, in the graph; the graph is then convolved to refine the features of objects and relations, and spatial sequence of the objects and relations is further modelled by Bi-LSTM; the attention-based object–relation feature fusion module aggregates the features of the objects and relations. SceneContextNet explicitly connects object detection, relationship modelling, graph construction, context reasoning, and interpretable prediction in one unified flow, unlike traditional module-stacking methods. We integrate explainability by visualising influential image regions, objects, and relationships through Grad-CAM, LIME, and graph-edge visualisation. Experimental results on ADE20K, Visual Genome and SUN RGB-D indicate that SceneContextNet can reach an accuracy of 98.62%, an F1-score of 98.51%, and an mAP of 98.30%, which are superior to those of CNN-, transformer-, and graph-based baselines. Further, cross-dataset evaluation shows strong generalisation across different scene layouts and object distributions. The proposed framework can be utilised in autonomous systems, intelligent surveillance, assistive vision, and smart environments where accurate and interpretable scene understanding is needed.

Discover ComputingVol. 29(1)
Koneru Lakshmaiah Education Foundation (IN)
Openalex Percentile: Top 14%
Multimodal Machine Learning Applications
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.