Intelligent Visual Prioritization for Retinal Prostheses via Context-Aware Object Ranking and Depth-Aware Phosphene Generation
Images from high-resolution cameras are mapped onto a sparse pattern of low spatial resolution and intensity in the retina, which limits visual perception in retinal prosthetic vision. When the entire scene is converted into phosphenes, it may allow unnecessary background information to be retained and may cause visual clutter, which may make it hard for prosthetic vision users to interpret the scene. In order to tackle this issue, this paper presents a context-, depth-, and user-preference-aware method for selecting the objects of interest in the generation of phosphene images. The proposed method does not show all the objects equally but learns to sort the objects according to their relevance to prosthetic vision. Manual annotation of a subset of COCO images was conducted where the most salient object was selected based on environment type, scene type, user mode, safety, navigation relevance, task importance, and distance. All of the candidate objects are described by full-scene visual features, object-crop features, handcrafted priority features, context embeddings, and monocular depth features. To predict object-level importance scores and identify the Top-1 and Top-4 important objects in unseen scenes, a hybrid deep learning model combining twin ResNet-18 backbones for scene and object feature extraction with embedding-based context encoding was trained. Priority maps and phosphene images were then created using the selected object masks and were depth-weighted. Two types of phosphene representations were also produced: Canny-edge-based and direct full images. The proposed framework is designed to suppress irrelevant background areas and improve important and closer objects in order to obtain a simplified and informative prosthetic-vision representation of the scene. The experimental evaluation, including Top-1 accuracy, Top-3 accuracy, mean reciprocal rank (MRR), and visual comparison, demonstrates the effectiveness of the proposed framework, achieving a Top-1 accuracy of 90.12%, a Top-3 accuracy of 97.45%, and an MRR of 0.9368. Furthermore, the proposed Canny-priority phosphene representation achieved an average human-participant recognition accuracy of approximately 86%. The proposed method offers a user-adaptive strategy for selecting and visualizing the information of a scene under the severe constraint of the bandwidth of retinal prosthetic vision.
Authors
- Muhammad Nawaz Khan (ORCID: https://orcid.org/0000-0002-6682-7049)
- Irshad Khalil (ORCID: https://orcid.org/0000-0002-8044-6779)
- Xinwei Li (ORCID: https://orcid.org/0000-0003-0911-5130)
- Faisal Rahman
Institutions
- Gachon University (KR)
- Liaocheng University (CN)
- Southwest Baptist University (US)
Publication Details
- Journal
- Biomimetics
- Published
- 2026-09-09
- DOI
- https://doi.org/10.3390/biomimetics11090649
- Primary Topic
- Neuroscience and Neural Engineering
- Type
- article
- Field-Weighted Citation Impact
- 0.00