A large-scale pipeline for LLM-assisted corpus annotation
Abstract As natural language corpora expand at an unprecedented rate, manual annotation remains a significant methodological bottleneck in corpus linguistic work. We address this challenge by presenting a scalable pipeline for automating grammatical annotation using large language models (LLMs). We demonstrate the pipeline’s accessibility and effectiveness through a diachronic case study of variation in the English evaluative consider construction ( consider X as / to be /Ø Y). We annotate 143,933 ‘consider’ concordance lines from the Corpus of Historical American English (COHA) via the OpenAI API in under 60 hours, achieving 98%+ accuracy on two sophisticated annotation procedures. Subsequent analysis reveals previously undocumented genre-specific trajectories of change, enabling us to advance new hypotheses about the relationship between register formality and competing pressures of morphosyntactic reduction and enhancement. Our results suggest that LLMs can perform a range of data preparation tasks at scale with minimal human intervention, unlocking substantive research questions previously beyond practical reach.
Authors
- Matti Marttinen Larsson (ORCID: https://orcid.org/0000-0002-6224-7872)
- Cameron Morin (ORCID: https://orcid.org/0000-0001-7079-449X)
Institutions
- Stockholm University (SE)
- Université Paris Cité (FR)
Publication Details
- Journal
- International Journal of Corpus Linguistics
- Published
- 2026-10-09
- DOI
- https://doi.org/10.1075/ijcl.25162.mor
- Primary Topic
- Natural Language Processing Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00