Metadata and annotation in historical text corpora: standards, challenges and best practices

Abstract The construction of historical text corpora presents distinctive challenges in terms of metadata curation, annotation practices and long-term accessibility. This article explores how standards-based approaches can enhance the discoverability, interoperability and sustainability of historical linguistic data. Emphasis is placed on aligning corpus design with the FAIR principles (Findable, Accessible, Interoperable, Reusable) (Wilkinson et al. 2016), which are increasingly important for research transparency and cross-platform integration in the digital humanities. After outlining the conceptual landscape of metadata and annotation in corpus linguistics, the article examines several existing projects that demonstrate the operationalisation of metadata standards across genres and temporal ranges. These include corpora that integrate CMDI profiles for complex resources (Paquot et al. 2024), and others that have transformed legacy metadata into machine-readable formats to facilitate data exchange and semantic enrichment (Fallucchi & De Luca 2020). Key issues addressed include orthographic variation, the limits of automated tagging tools for historical data (Pettersson & Megyesi 2018) and the need to balance standardisation with the preservation of linguistic idiosyncrasies. The article explores a case study of the MetaLing Corpus , a one-million-token historical corpus of English metalanguage from 1500 to 1700 (Andreani & Russo 2026). Practical constraints and editorial decisions are discussed. A minimalist encoding strategy was adopted following compatibility issues with TEI and Sketch Engine (Kilgarriff et al. 2004; Kilgarriff et al. 2014), resulting in a plain-text corpus linked to metadata managed through Omeka using Dublin Core (Caplan 2003). This study contributes to ongoing discussions on metadata sustainability, annotation design and best practices for corpus construction in underrepresented linguistic domains. It advocates for adaptable, transparent approaches that foster cross-disciplinary collaboration and position corpora as evolving infrastructures within broader digital ecosystems.

Authors

Institutions

Publication Details

Journal
English Language and Linguistics
Published
2026-09-29
DOI
https://doi.org/10.1017/s1360674326100847
Primary Topic
Diverse Musicological Studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Metadata and annotation in historical text corpora: standards, challenges and best practices

Dmitriy Eduardovich Russo
English Language and Linguistics
Diverse Musicological Studies
article

Metadata and annotation in historical text corpora: standards, challenges and best practices

Dmitriy Eduardovich Russo
article en

Abstract

Abstract The construction of historical text corpora presents distinctive challenges in terms of metadata curation, annotation practices and long-term accessibility. This article explores how standards-based approaches can enhance the discoverability, interoperability and sustainability of historical linguistic data. Emphasis is placed on aligning corpus design with the FAIR principles (Findable, Accessible, Interoperable, Reusable) (Wilkinson et al. 2016), which are increasingly important for research transparency and cross-platform integration in the digital humanities. After outlining the conceptual landscape of metadata and annotation in corpus linguistics, the article examines several existing projects that demonstrate the operationalisation of metadata standards across genres and temporal ranges. These include corpora that integrate CMDI profiles for complex resources (Paquot et al. 2024), and others that have transformed legacy metadata into machine-readable formats to facilitate data exchange and semantic enrichment (Fallucchi & De Luca 2020). Key issues addressed include orthographic variation, the limits of automated tagging tools for historical data (Pettersson & Megyesi 2018) and the need to balance standardisation with the preservation of linguistic idiosyncrasies. The article explores a case study of the MetaLing Corpus , a one-million-token historical corpus of English metalanguage from 1500 to 1700 (Andreani & Russo 2026). Practical constraints and editorial decisions are discussed. A minimalist encoding strategy was adopted following compatibility issues with TEI and Sketch Engine (Kilgarriff et al. 2004; Kilgarriff et al. 2014), resulting in a plain-text corpus linked to metadata managed through Omeka using Dublin Core (Caplan 2003). This study contributes to ongoing discussions on metadata sustainability, annotation design and best practices for corpus construction in underrepresented linguistic domains. It advocates for adaptable, transparent approaches that foster cross-disciplinary collaboration and position corpora as evolving infrastructures within broader digital ecosystems.

English Language and Linguistics
University of Insubria (IT)
Industry, innovation and infrastructure
Openalex Percentile: Top 2%
Diverse Musicological Studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.