Linguistic Asymmetry in the CLARIN Virtual Language Observatory
Presentation by researcher Elena Battaner Moro for the CLARIN Annual Conference 2026 (Brighton) on Linguistic Asymmetry in the CLARIN VLO. The CLARIN Virtual Language Observatory (VLO) brings together metadata from distributed repositories and allows users to discover language resources through a shared interface. This paper asks how consistently the VLO languageCode facet supports discovery by language. Using a complete snapshot of the field, we examine the distribution of code-based and textual language values and conduct a conservative ISO 639 audit of the textual labels. The results show broad but uneven representation, frequent indeterminate language values, and confirmed cases in which the same language is exposed through both an ISO code and a separate textual label, creating potential retrieval fragmentation. Multilingual coverage should therefore be assessed not only by the number of languages present, but also by how language information is represented and exposed for search. We recommend lightweight reconciliation for unambiguous cases, clearer provider guidance, and repeatable monitoring of linguistic representation.
Authors
- Elena Battaner Moro (ORCID: https://orcid.org/0000-0002-5521-6445)
Institutions
- Universidad Rey Juan Carlos (ES)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-01
- DOI
- https://doi.org/10.5281/zenodo.23080356
- Primary Topic
- Natural Language Processing Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00