SceneContextNet: an integrated multi-stage deep learning framework for context-aware scene understanding using multi-object detection and spatial relationship modelling
Deep learning, especially CNN- and transformer-based approaches, has significantly impacted computer vision techniques such as object detection, segmentation, and scene classification. Despite the many successful scene understanding pipelines currently in use, most still rely primarily on global image features or object-level predictions and have little capacity to explicitly model object relationships, spatial arrangements, and structured contextual dependencies. This limits their performance in complex scenes with multiple objects in different semantic layouts. To address this limitation, this paper introduces SceneContextNet. This relationship-guided scene understanding framework transforms detected objects into a structured scene graph and reasons over object–relation representations with context. The novel idea of the proposed approach is to employ a spatial–semantic graph-sequence reasoning strategy, which encodes the object features and pairwise relationship descriptors as nodes and edges, respectively, in the graph; the graph is then convolved to refine the features of objects and relations, and spatial sequence of the objects and relations is further modelled by Bi-LSTM; the attention-based object–relation feature fusion module aggregates the features of the objects and relations. SceneContextNet explicitly connects object detection, relationship modelling, graph construction, context reasoning, and interpretable prediction in one unified flow, unlike traditional module-stacking methods. We integrate explainability by visualising influential image regions, objects, and relationships through Grad-CAM, LIME, and graph-edge visualisation. Experimental results on ADE20K, Visual Genome and SUN RGB-D indicate that SceneContextNet can reach an accuracy of 98.62%, an F1-score of 98.51%, and an mAP of 98.30%, which are superior to those of CNN-, transformer-, and graph-based baselines. Further, cross-dataset evaluation shows strong generalisation across different scene layouts and object distributions. The proposed framework can be utilised in autonomous systems, intelligent surveillance, assistive vision, and smart environments where accurate and interpretable scene understanding is needed.
Authors
- Venubabu Rachapudi
- M. Srividya
Institutions
- Koneru Lakshmaiah Education Foundation (IN)
Publication Details
- Journal
- Discover Computing
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1007/s10791-026-10615-x
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00