SAVPEF-Net: Symmetric Semantic–Visual Prompt Fusion for Fine-Grained Open-World Object Detection in Intelligent Transportation
Open-world object detection is a key visual support technology for smart-city transportation. Most existing methods are trained on closed-set datasets and recognize only predefined traffic categories. For unknown traffic participants and sudden obstacles in complex scenes, they suffer severe false/missed detection and cannot adapt to dynamic transportation. Mainstream open-vocabulary methods rely on complex cross-modal fusion, and their inference overhead increases sharply with vocabulary scale, making it difficult to meet real-time traffic edge requirements. They also face bottlenecks: insufficient symmetric adaptability between fixed text embeddings and traffic scenarios, low unknown-object accuracy, and difficult category expansion. This paper proposes a fine-grained open-world object detection network for intelligent transportation enhanced by semantically activated visual prompts, realizing accurate and efficient recognition via symmetric multimodal feature alignment, visual prompt adaptation, and symmetric unknown object perception. First, we propose an Adaptive Decision Text Encoding Module, which achieves lightweight end-to-end symmetric text–visual alignment through low-rank adaptation and reparameterization, eliminating cross-modal fusion redundancy while retaining CLIP (Contrastive Language-Image Pre-training) generalization. Second, a Semantically Activated Visual Prompt Encoding Module realizes low-dimensional efficient representation of traffic visual cues through decoupled symmetric dual branches of semantics and activation, supplementing text-prompt scene adaptability in fine-grained detection. Then, a Wildcard Lazy Fusion Module combines symmetric dual-wildcard self-supervised learning and lazy region contrast to accurately recognize known objects and effectively detect unknown obstacles and traffic participants, while supporting dynamic category expansion. Finally, a Dual-Head Collaborative Detection Module jointly optimizes object localization and classification through a cross-head symmetric collaborative matching strategy, achieving end-to-end efficient inference without NMS (Non-Maximum Suppression). Extensive experiments show that compared with advanced open-world object detection networks, the proposed method achieves 2.1% and 2.6% absolute improvements in unknown object recall on DOTA-v1.0 and UCAS-AOD respectively relative to the RT-DETR (Rotated Detection Transformer) baseline, while delivering a 1.0% gain in rotation mean average precision (mAP) for known categories and sustaining 31.4 FPS on a single NVIDIA RTX 3090 GPU.
Authors
- Donglin Jing (ORCID: https://orcid.org/0000-0003-3021-5371)
- Zhengbiao Jing (ORCID: https://orcid.org/0009-0002-5693-7946)
- Douping Bai
Publication Details
- Journal
- Algorithms
- Published
- 2026-10-04
- DOI
- https://doi.org/10.3390/a19100850
- Primary Topic
- Advanced Neural Network Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00