Linguistic Asymmetry in the CLARIN Virtual Language Observatory

Presentation by researcher Elena Battaner Moro for the CLARIN Annual Conference 2026 (Brighton) on Linguistic Asymmetry in the CLARIN VLO. The CLARIN Virtual Language Observatory (VLO) brings together metadata from distributed repositories and allows users to discover language resources through a shared interface. This paper asks how consistently the VLO languageCode facet supports discovery by language. Using a complete snapshot of the field, we examine the distribution of code-based and textual language values and conduct a conservative ISO 639 audit of the textual labels. The results show broad but uneven representation, frequent indeterminate language values, and confirmed cases in which the same language is exposed through both an ISO code and a separate textual label, creating potential retrieval fragmentation. Multilingual coverage should therefore be assessed not only by the number of languages present, but also by how language information is represented and exposed for search. We recommend lightweight reconciliation for unambiguous cases, clearer provider guidance, and repeatable monitoring of linguistic representation.

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-01
DOI
https://doi.org/10.5281/zenodo.23080356
Primary Topic
Natural Language Processing Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Linguistic Asymmetry in the CLARIN Virtual Language Observatory

Elena Battaner Moro
Zenodo (CERN European Organization for Nuclear Research)
Natural Language Processing Techniques
article

Linguistic Asymmetry in the CLARIN Virtual Language Observatory

Elena Battaner Moro
article en

Abstract

Presentation by researcher Elena Battaner Moro for the CLARIN Annual Conference 2026 (Brighton) on Linguistic Asymmetry in the CLARIN VLO. The CLARIN Virtual Language Observatory (VLO) brings together metadata from distributed repositories and allows users to discover language resources through a shared interface. This paper asks how consistently the VLO languageCode facet supports discovery by language. Using a complete snapshot of the field, we examine the distribution of code-based and textual language values and conduct a conservative ISO 639 audit of the textual labels. The results show broad but uneven representation, frequent indeterminate language values, and confirmed cases in which the same language is exposed through both an ISO code and a separate textual label, creating potential retrieval fragmentation. Multilingual coverage should therefore be assessed not only by the number of languages present, but also by how language information is represented and exposed for search. We recommend lightweight reconciliation for unambiguous cases, clearer provider guidance, and repeatable monitoring of linguistic representation.

Zenodo (CERN European Organization for Nuclear Research)
Universidad Rey Juan Carlos (ES)
Quality Education
Openalex Percentile: Top 9%
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Linguistic Asymmetry in the CLARIN Virtual Language Observatory — Elena Battaner Moro · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS