Adaptive speech-to-image translation for impaired speech using parameter-efficient ASR and diffusion-based image generation
Abstract Impaired Speech Continues to Pose Challenges for the Automatic Processing of Speech and Multimodal Interaction between Humans and Computers systems caused by distortion of articulation and involuntary repetition. Converting impaired speech into sight allows a complementary system of communicating for individuals who have speech disorders; especially when the textual input is often the only method of expression difficult to interpret. This study proposes a cascaded speech-to-image translation framework which integrates adaptive speech recognition using the diffusion-based image generating. The speech recognition part is improved based on Silero-based voice activity detection, parameter-efficient fine-tuning using AdaLoRA and shallow fusion using a large language model to improve under dysarthric conditions the robustness of the transcription; The refined textual output is then used to condition a latent diffusion model optimized for controlled image synthesis. Experiments were conducted on the publicly available TORGO dysarthric speech corpus. The proposed configuration achieved a word error rate (WER) of 0.125 and a character error rate (CER) of 0.050, substantially outperforming the baseline Whisper models. Semantic correspondence between generated images and their conditioning text prompts are quantified using CLIPScore with 30.18 a score of 30.18 for the proposed system. indicating spread of multimodal consistency compared to non-adaptive baselines. Qualitative results further show that the framework maintains important semantic features across diverse categories of prompts such as natural scenes, human portraits and action-oriented layouts. These results indicated that adaptation of speech recognition plays a critical role in stabilizing downstream generating visuals and supporting the feasibility of speech-to-image translation as an assistive communication modality for impaired speech.
Authors
- A. Bazila Banu
- Philo Sumi
Institutions
- KPR Institute of Engineering and Technology (IN)
Publication Details
- Journal
- Scientific Reports
- Published
- 2026-09-24
- DOI
- https://doi.org/10.1038/s41598-026-62752-4
- Primary Topic
- Voice and Speech Disorders
- Type
- article
- Field-Weighted Citation Impact
- 0.00