The Azerbaijani Speech-to-Text Landscape: Data, Models, and the Gap Between Claimed and Demonstrated Support

Azerbaijani is spoken by roughly ten million people and listed as supported by thirteen commercial speech-recognition services. We surveyed what actually exists behind those claims. Only one publishes a numeric Azerbaijani accuracy figure, and it appears on a marketing page rather than in its documentation; two others publish coarse accuracy bands that disagree by a full tier on undisclosed test sets. Exactly one vendor documents a purpose-built Azerbaijani acoustic model, and everything else in the market is a rebranded Whisper or MMS. Twelve papers have Azerbaijani speech recognition as their primary subject, none of them on arXiv, in the ACL Anthology, or at Interspeech or ICASSP, and not one evaluates on a public test set. The best published result on the only shared benchmark is 19.8% word error rate; the best known result, 13.17%, appears in a model card with no paper behind it, so the state of the art for this language is literally unpublished. The public data is almost entirely read and broadcast speech: we find zero hours of public spontaneous or telephone-band Azerbaijani, while every natively narrowband open toolkit -- Vosk, Kaldi, NeMo, Coqui -- has no Azerbaijani model at all. Common Voice holds 0.65 validated hours against Turkish's 128, a gap of roughly 260 to 1 between two mutually intelligible Oghuz languages. We further documenttwo previously unreported defects in the FLEURS Azerbaijani transcripts that corrupt character-level metrics on the field's only shared benchmark. We release the survey's evidence base as a versioned, machine-readable roster of 44 artifacts in which every entry is marked as measured, documented, or merely claimed.

Authors

Publication Details

Journal
arXiv (Cornell University)
Published
2026-09-14
DOI
https://doi.org/10.5281/zenodo.22742915
Primary Topic
Natural Language Processing Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

The Azerbaijani Speech-to-Text Landscape: Data, Models, and the Gap Between Claimed and Demonstrated Support

O. Ibrahimzade
arXiv (Cornell University)
Natural Language Processing Techniques
preprint

The Azerbaijani Speech-to-Text Landscape: Data, Models, and the Gap Between Claimed and Demonstrated Support

O. Ibrahimzade
preprint en

Abstract

Azerbaijani is spoken by roughly ten million people and listed as supported by thirteen commercial speech-recognition services. We surveyed what actually exists behind those claims. Only one publishes a numeric Azerbaijani accuracy figure, and it appears on a marketing page rather than in its documentation; two others publish coarse accuracy bands that disagree by a full tier on undisclosed test sets. Exactly one vendor documents a purpose-built Azerbaijani acoustic model, and everything else in the market is a rebranded Whisper or MMS. Twelve papers have Azerbaijani speech recognition as their primary subject, none of them on arXiv, in the ACL Anthology, or at Interspeech or ICASSP, and not one evaluates on a public test set. The best published result on the only shared benchmark is 19.8% word error rate; the best known result, 13.17%, appears in a model card with no paper behind it, so the state of the art for this language is literally unpublished. The public data is almost entirely read and broadcast speech: we find zero hours of public spontaneous or telephone-band Azerbaijani, while every natively narrowband open toolkit -- Vosk, Kaldi, NeMo, Coqui -- has no Azerbaijani model at all. Common Voice holds 0.65 validated hours against Turkish's 128, a gap of roughly 260 to 1 between two mutually intelligible Oghuz languages. We further documenttwo previously unreported defects in the FLEURS Azerbaijani transcripts that corrupt character-level metrics on the field's only shared benchmark. We release the survey's evidence base as a versioned, machine-readable roster of 44 artifacts in which every entry is marked as measured, documented, or merely claimed.

arXiv (Cornell University)
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

The Azerbaijani Speech-to-Text Landscape: Data, Models, and the Gap Between Claimed and Demonstrated Support — O. Ibrahimzade · arXiv (Cornell University) (2026) | TGRS Research Map | TGRS