Quality gated multimodal fusion improves scene construction and interaction experience in augmented reality digital media
Augmented reality now carries a great deal of digital media production, yet its scene construction still suffers from registration drift, asynchronous sensor streams, unnatural interaction and excessive cognitive load. We address these coupled problems with an integrated framework. Hardware-triggered acquisition and timestamping disciplined by the Precision Time Protocol feed a quality-gated cross-modal attention model whose fusion weights follow per-modality reliability rather than fixed priors. A dynamic scene construction algorithm then couples semantic matching, layout optimization by covariance matrix adaptation and harmonic relighting, so that virtual content stays geometrically and photometrically consistent with the room around it, while an adaptive feedback policy driven by the decoder’s uncertainty estimates tunes interaction cues to the demand of the current task. On a corpus of 184 controlled sequences together with public benchmark subsets, running on a HoloLens 2 platform, the system reaches 91.4% fusion accuracy at a median end-to-end latency of 16.8 ms with a 95th percentile of 18.9 ms, and its robustness margin widens when two modalities degrade at once. A within-subjects user study (n = 50) recorded reliable gains in immersion, presence, usability, satisfaction and cognitive load, with an exploratory analysis suggesting that first-time users benefit more than experienced ones. The approach offers a coherent path toward AR digital media that remains usable beyond enthusiast audiences, although outdoor scalability, hardware dependence and long-session fatigue all remain open.
Authors
- Jundan Wang
- Yifan Zhang
Institutions
- Hoseo University (KR)
- Huizhou University (CN)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-06
- DOI
- https://doi.org/10.1038/s41598-026-69282-z
- Primary Topic
- Augmented Reality Applications
- Type
- article
- Field-Weighted Citation Impact
- 0.00