Adapting Kokoro-82M to Bengali: A Teacher-Forced Fine-Tuning Recipe for a Text-to-Speech Model Released Without Training Code

Kokoro-82M is a compact, high-quality text-to-speech (TTS) model whose weights are public but whose training code is not. Its released voices cover a handful of languages; Bengali is not one of them. This report describes a complete recipe for adding a new language to Kokoro-82M using only the published inference package and an external forced aligner. We reconstruct the StyleTTS2-style teacher-forced training objective on top of the unmodified model modules, obtain phoneme durations from Meta's MMS forced aligner via romanised word alignment, map espeak-ng phonemes onto Kokoro's fixed 178-symbol vocabulary, and train on 18.3 hours of single-speaker studio Bengali from the IIT Madras IndicTTS corpus. A first attempt that drove the vocoder with the model's own predicted pitch and energy produced muffled, harmonic-poor speech that did not improve between 1,000 and 2,000 steps. Switching the vocoder input to ground-truth pitch and energy during training, as StyleTTS2 does, restored clear harmonic structure at 1,000 steps and raised output loudness by 5.5 dB toward the reference. Training is ongoing; we report the pipeline, the ablation, and objective measurements, and outline the perceptual evaluation planned once training completes.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-09-16
DOI
https://doi.org/10.5281/zenodo.22789547
Primary Topic
Speech Recognition and Synthesis
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Adapting Kokoro-82M to Bengali: A Teacher-Forced Fine-Tuning Recipe for a Text-to-Speech Model Released Without Training Code

Sunny Kumar
Zenodo (CERN European Organization for Nuclear Research)
Speech Recognition and Synthesis
preprint

Adapting Kokoro-82M to Bengali: A Teacher-Forced Fine-Tuning Recipe for a Text-to-Speech Model Released Without Training Code

Sunny Kumar
preprint en

Abstract

Kokoro-82M is a compact, high-quality text-to-speech (TTS) model whose weights are public but whose training code is not. Its released voices cover a handful of languages; Bengali is not one of them. This report describes a complete recipe for adding a new language to Kokoro-82M using only the published inference package and an external forced aligner. We reconstruct the StyleTTS2-style teacher-forced training objective on top of the unmodified model modules, obtain phoneme durations from Meta's MMS forced aligner via romanised word alignment, map espeak-ng phonemes onto Kokoro's fixed 178-symbol vocabulary, and train on 18.3 hours of single-speaker studio Bengali from the IIT Madras IndicTTS corpus. A first attempt that drove the vocoder with the model's own predicted pitch and energy produced muffled, harmonic-poor speech that did not improve between 1,000 and 2,000 steps. Switching the vocoder input to ground-truth pitch and energy during training, as StyleTTS2 does, restored clear harmonic structure at 1,000 steps and raised output loudness by 5.5 dB toward the reference. Training is ongoing; we report the pipeline, the ablation, and objective measurements, and outline the perceptual evaluation planned once training completes.

Zenodo (CERN European Organization for Nuclear Research)
Indian Institute of Technology Patna (IN)
Quality Education
Speech Recognition and Synthesis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Adapting Kokoro-82M to Bengali: A Teacher-Forced Fine-Tuning Recipe for a Text-to-Speech Model Released Without Training Code — Sunny Kumar · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS