Context-Aware visual scene analysis and adaptation for intelligent english tutoring systems
Abstract The integration of visual context into English language teaching is essential for fostering “situated cognition” in learners. However, current Intelligent English Tutoring Systems often lack the ability to interpret complex visual environments on a massive scale. This paper introduces a visual scene categorization method designed to support context-sensitive English education. Our approach uses BING objectness measures to identify object-centric candidate regions that can serve as potential visual cues for vocabulary and scene understanding. For an input image, BING ranks candidate regions by objectness score, and the top- L retained regions form an objectness-guided selected-patch set . A shared CNN extracts descriptors from the selected patches, which are statistically aggregated and modeled by a Gaussian Mixture Model to categorize learning scenarios such as airports, supermarkets, and classrooms. This procedure is an automatic region-selection and representation-learning pipeline; it does not reconstruct a human visual scanpath. Experiments conducted on a massive-scale scenery dataset demonstrate that the method identifies meaningful educational contexts for context-aware tutoring.
Authors
- Yue Yu
- Huihua Wu
Institutions
- Hainan Agricultural School (CN)
- Jinhua University of Vocational Technology (CN)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-19
- DOI
- https://doi.org/10.1038/s41598-026-70908-5
- Primary Topic
- Multimodal Machine Learning Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00