Learning Finnish inflectional classes: experiments with the discriminative lexicon model
Abstract This study shows that the Discriminative Lexicon Model (DLM) can learn Finnish nominal inflection from paired form and meaning embeddings, without being given access to stems, exponents, inflectional features, and information about inflectional class. It also shows that the model is productive. It generalizes to novel words, and does so with greater precision for more productive inflectional classes. This does not imply that productivity reduces simply to the absence of irregularity. Rather, generalization is made possible thanks to substantial isomorphies that exist between the space of form embeddings and the space of meaning embeddings. Usage-based DLM models that take token frequency into account have lower type accuracy: practice makes perfect, but with less practice, learning becomes increasingly problematic. Nevertheless, token-based DLMs perform well token-wise; they generalize to novel words, and reflect the differences in productivity of inflectional classes even better than type-based models. It is also shown that DLM comprehension models with form embeddings based on 4-grams outperform models with form embeddings based on 3-grams. This is due to 4-grams better covering the partial combinatorics of exponents and corresponding to more informative regions in the semantic embedding space.
Authors
- Alexandre Nikolaev (ORCID: https://orcid.org/0000-0001-8634-5947)
- Yu‐Ying Chuang (ORCID: https://orcid.org/0000-0002-2733-2748)
- R. Harald Baayen (ORCID: https://orcid.org/0000-0003-3178-3944)
Institutions
- National Taiwan Normal University (TW)
- University of Eastern Finland (FI)
- University of Tübingen (DE)
Publication Details
- Journal
- Morphology
- Published
- 2026-09-17
- DOI
- https://doi.org/10.1007/s11525-026-09470-9
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- Itä-Suomen Yliopisto
- Kuopion Yliopistollinen Sairaala