FROM WORDS TO WORLDS: A MULTIMODAL ANNOTATION FRAMEWORK FOR POST-TANG CHINESE LYRIC

This paper presents a methodological framework for the digital curation of the post-Tang Chinese lyric — the ci, qu, and sanqu composed for musical realization between the late Tang and the end of the Qing (ninth century to 1911). The framework coordinates three mutually constraining annotation layers: a textual layer containing the canonical poem, variants, and the commentarial tradition; a multimodal layer linking the text to its notated music — chiefly the jianzipu (减字谱) tablature of the qin-song repertory, alongside the suzipu pitch notation of the sung ci — and to its calligraphic carriers (ink tracings, rubbings, engraved transcriptions); and a Six-W provenance layer capturing who wrote, addressed, and recited the poem, what it sets out to do, when and where it was composed, why it was occasioned, and how it was transmitted. Situated against contemporary annotation and provenance models (TEI, W3C PROV-O, FRBR, CIDOC-CRM, Dublin Core, and the W7 model), the paper articulates an AI-assisted annotation pipeline — NLP for the textual layer, computer vision and handwriting recognition for the calligraphic layer, optical music recognition for the score layer — under a human-in-the-loop protocol that preserves the philological prerogative of interpretive judgment. The argument is illustrated by one fully worked case drawn from that dataset, the qin song Gu Yuan (古怨) in Jiang Kui’s Baishidaoren Gequ, whose surviving editions support a documented evidential gap that the schema is designed to carry rather than resolve. The framework is not a proposal awaiting implementation: the pipeline described here has been built and run across all three annotation layers, and the resulting dataset comprises approximately 1.3 million post-Tang lyric compositions, each carrying a Six-W provenance record, with the full three-layer treatment applied wherever a non-textual witness survives. The dataset is not yet publicly released; the schema, the pipeline, and the epistemic principles are set out here in full so that the data may be assessed independently of the release schedule.

Authors

Institutions

Publication Details

Journal
Journal of Humanities and AI
Published
2026-09-30
DOI
https://doi.org/10.66532/jhai.2026.0021
Primary Topic
Music and Audio Processing
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

FROM WORDS TO WORLDS: A MULTIMODAL ANNOTATION FRAMEWORK FOR POST-TANG CHINESE LYRIC

Yao Song
Journal of Humanities and AI
Music and Audio Processing
article

FROM WORDS TO WORLDS: A MULTIMODAL ANNOTATION FRAMEWORK FOR POST-TANG CHINESE LYRIC

Yao Song
article en

Abstract

This paper presents a methodological framework for the digital curation of the post-Tang Chinese lyric — the ci, qu, and sanqu composed for musical realization between the late Tang and the end of the Qing (ninth century to 1911). The framework coordinates three mutually constraining annotation layers: a textual layer containing the canonical poem, variants, and the commentarial tradition; a multimodal layer linking the text to its notated music — chiefly the jianzipu (减字谱) tablature of the qin-song repertory, alongside the suzipu pitch notation of the sung ci — and to its calligraphic carriers (ink tracings, rubbings, engraved transcriptions); and a Six-W provenance layer capturing who wrote, addressed, and recited the poem, what it sets out to do, when and where it was composed, why it was occasioned, and how it was transmitted. Situated against contemporary annotation and provenance models (TEI, W3C PROV-O, FRBR, CIDOC-CRM, Dublin Core, and the W7 model), the paper articulates an AI-assisted annotation pipeline — NLP for the textual layer, computer vision and handwriting recognition for the calligraphic layer, optical music recognition for the score layer — under a human-in-the-loop protocol that preserves the philological prerogative of interpretive judgment. The argument is illustrated by one fully worked case drawn from that dataset, the qin song Gu Yuan (古怨) in Jiang Kui’s Baishidaoren Gequ, whose surviving editions support a documented evidential gap that the schema is designed to carry rather than resolve. The framework is not a proposal awaiting implementation: the pipeline described here has been built and run across all three annotation layers, and the resulting dataset comprises approximately 1.3 million post-Tang lyric compositions, each carrying a Six-W provenance record, with the full three-layer treatment applied wherever a non-textual witness survives. The dataset is not yet publicly released; the schema, the pipeline, and the epistemic principles are set out here in full so that the data may be assessed independently of the release schedule.

Journal of Humanities and AIVol. 1(3)
Hong Kong Baptist University (HK)
Openalex Percentile: Top 11%
Music and Audio Processing
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

FROM WORDS TO WORLDS: A MULTIMODAL ANNOTATION FRAMEWORK FOR POST-TANG CHINESE LYRIC — Yao Song · Journal of Humanities and AI (2026) | TGRS Research Map | TGRS