The Azerbaijani Speech-to-Text Landscape: Data, Models, and the Gap Between Claimed and Demonstrated Support
Azerbaijani is spoken by roughly ten million people and listed as supported by thirteen commercial speech-recognition services. We surveyed what actually exists behind those claims. Only one publishes a numeric Azerbaijani accuracy figure, and it appears on a marketing page rather than in its documentation; two others publish coarse accuracy bands that disagree by a full tier on undisclosed test sets. Exactly one vendor documents a purpose-built Azerbaijani acoustic model, and everything else in the market is a rebranded Whisper or MMS. Twelve papers have Azerbaijani speech recognition as their primary subject, none of them on arXiv, in the ACL Anthology, or at Interspeech or ICASSP, and not one evaluates on a public test set. The best published result on the only shared benchmark is 19.8% word error rate; the best known result, 13.17%, appears in a model card with no paper behind it, so the state of the art for this language is literally unpublished. The public data is almost entirely read and broadcast speech: we find zero hours of public spontaneous or telephone-band Azerbaijani, while every natively narrowband open toolkit -- Vosk, Kaldi, NeMo, Coqui -- has no Azerbaijani model at all. Common Voice holds 0.65 validated hours against Turkish's 128, a gap of roughly 260 to 1 between two mutually intelligible Oghuz languages. We further documenttwo previously unreported defects in the FLEURS Azerbaijani transcripts that corrupt character-level metrics on the field's only shared benchmark. We release the survey's evidence base as a versioned, machine-readable roster of 44 artifacts in which every entry is marked as measured, documented, or merely claimed.
Authors
- O. Ibrahimzade
Publication Details
- Journal
- arXiv (Cornell University)
- Published
- 2026-09-14
- DOI
- https://doi.org/10.5281/zenodo.22742915
- Primary Topic
- Natural Language Processing Techniques
- Type
- preprint