Adaptive speech-to-image translation for impaired speech using parameter-efficient ASR and diffusion-based image generation

Abstract Impaired Speech Continues to Pose Challenges for the Automatic Processing of Speech and Multimodal Interaction between Humans and Computers systems caused by distortion of articulation and involuntary repetition. Converting impaired speech into sight allows a complementary system of communicating for individuals who have speech disorders; especially when the textual input is often the only method of expression difficult to interpret. This study proposes a cascaded speech-to-image translation framework which integrates adaptive speech recognition using the diffusion-based image generating. The speech recognition part is improved based on Silero-based voice activity detection, parameter-efficient fine-tuning using AdaLoRA and shallow fusion using a large language model to improve under dysarthric conditions the robustness of the transcription; The refined textual output is then used to condition a latent diffusion model optimized for controlled image synthesis. Experiments were conducted on the publicly available TORGO dysarthric speech corpus. The proposed configuration achieved a word error rate (WER) of 0.125 and a character error rate (CER) of 0.050, substantially outperforming the baseline Whisper models. Semantic correspondence between generated images and their conditioning text prompts are quantified using CLIPScore with 30.18 a score of 30.18 for the proposed system. indicating spread of multimodal consistency compared to non-adaptive baselines. Qualitative results further show that the framework maintains important semantic features across diverse categories of prompts such as natural scenes, human portraits and action-oriented layouts. These results indicated that adaptation of speech recognition plays a critical role in stabilizing downstream generating visuals and supporting the feasibility of speech-to-image translation as an assistive communication modality for impaired speech.

Authors

Institutions

Publication Details

Journal
Scientific Reports
Published
2026-09-24
DOI
https://doi.org/10.1038/s41598-026-62752-4
Primary Topic
Voice and Speech Disorders
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Adaptive speech-to-image translation for impaired speech using parameter-efficient ASR and diffusion-based image generation

A. Bazila Banu, Philo Sumi
Scientific Reports
Voice and Speech Disorders
article

Adaptive speech-to-image translation for impaired speech using parameter-efficient ASR and diffusion-based image generation

A. Bazila Banu, Philo Sumi
article en

Abstract

Abstract Impaired Speech Continues to Pose Challenges for the Automatic Processing of Speech and Multimodal Interaction between Humans and Computers systems caused by distortion of articulation and involuntary repetition. Converting impaired speech into sight allows a complementary system of communicating for individuals who have speech disorders; especially when the textual input is often the only method of expression difficult to interpret. This study proposes a cascaded speech-to-image translation framework which integrates adaptive speech recognition using the diffusion-based image generating. The speech recognition part is improved based on Silero-based voice activity detection, parameter-efficient fine-tuning using AdaLoRA and shallow fusion using a large language model to improve under dysarthric conditions the robustness of the transcription; The refined textual output is then used to condition a latent diffusion model optimized for controlled image synthesis. Experiments were conducted on the publicly available TORGO dysarthric speech corpus. The proposed configuration achieved a word error rate (WER) of 0.125 and a character error rate (CER) of 0.050, substantially outperforming the baseline Whisper models. Semantic correspondence between generated images and their conditioning text prompts are quantified using CLIPScore with 30.18 a score of 30.18 for the proposed system. indicating spread of multimodal consistency compared to non-adaptive baselines. Qualitative results further show that the framework maintains important semantic features across diverse categories of prompts such as natural scenes, human portraits and action-oriented layouts. These results indicated that adaptation of speech recognition plays a critical role in stabilizing downstream generating visuals and supporting the feasibility of speech-to-image translation as an assistive communication modality for impaired speech.

Scientific Reports
KPR Institute of Engineering and Technology (IN)
Openalex Percentile: Top 11%
Voice and Speech Disorders
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Adaptive speech-to-image translation for impaired speech using parameter-efficient ASR and diffusion-based image generation — A. Bazila Banu, Philo Sumi · Scientific Reports (2026) | TGRS Research Map | TGRS