A large-scale pipeline for LLM-assisted corpus annotation

Abstract As natural language corpora expand at an unprecedented rate, manual annotation remains a significant methodological bottleneck in corpus linguistic work. We address this challenge by presenting a scalable pipeline for automating grammatical annotation using large language models (LLMs). We demonstrate the pipeline’s accessibility and effectiveness through a diachronic case study of variation in the English evaluative consider construction ( consider X as / to be /Ø Y). We annotate 143,933 ‘consider’ concordance lines from the Corpus of Historical American English (COHA) via the OpenAI API in under 60 hours, achieving 98%+ accuracy on two sophisticated annotation procedures. Subsequent analysis reveals previously undocumented genre-specific trajectories of change, enabling us to advance new hypotheses about the relationship between register formality and competing pressures of morphosyntactic reduction and enhancement. Our results suggest that LLMs can perform a range of data preparation tasks at scale with minimal human intervention, unlocking substantive research questions previously beyond practical reach.

Authors

Institutions

Publication Details

Journal
International Journal of Corpus Linguistics
Published
2026-10-09
DOI
https://doi.org/10.1075/ijcl.25162.mor
Primary Topic
Natural Language Processing Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

A large-scale pipeline for LLM-assisted corpus annotation

Matti Marttinen Larsson, Cameron Morin
International Journal of Corpus Linguistics
Natural Language Processing Techniques
article

A large-scale pipeline for LLM-assisted corpus annotation

Matti Marttinen Larsson, Cameron Morin
article en

Abstract

Abstract As natural language corpora expand at an unprecedented rate, manual annotation remains a significant methodological bottleneck in corpus linguistic work. We address this challenge by presenting a scalable pipeline for automating grammatical annotation using large language models (LLMs). We demonstrate the pipeline’s accessibility and effectiveness through a diachronic case study of variation in the English evaluative consider construction ( consider X as / to be /Ø Y). We annotate 143,933 ‘consider’ concordance lines from the Corpus of Historical American English (COHA) via the OpenAI API in under 60 hours, achieving 98%+ accuracy on two sophisticated annotation procedures. Subsequent analysis reveals previously undocumented genre-specific trajectories of change, enabling us to advance new hypotheses about the relationship between register formality and competing pressures of morphosyntactic reduction and enhancement. Our results suggest that LLMs can perform a range of data preparation tasks at scale with minimal human intervention, unlocking substantive research questions previously beyond practical reach.

International Journal of Corpus Linguistics
Stockholm University (SE), Université Paris Cité (FR)
Openalex Percentile: Top 13%
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

A large-scale pipeline for LLM-assisted corpus annotation — Matti Marttinen Larsson, Cameron Morin · International Journal of Corpus Linguistics (2026) | TGRS Research Map | TGRS