Automatic Annotation and Analysis of Oral History using LLMs: An Empirical Study of Densho Digital Collection

Oral histories serve as vital records of lived experience, particularly within communities affected by systemic injustice and historical erasure. Effective and efficient annotation and analysis of oral history archives can promote access and use of the oral histories. However, large scale analysis of these archives remains limited due to their unstructured format, emotional complexity, and the high cost of manual annotation. This paper presents a scalable LLM-based framework to automate semantic and sentiment annotation for oral history archives, with a focus on Japanese American Incarceration Oral History (JAIOH). Using large language models (LLMs), this study seeks to construct a high-quality dataset, systematically evaluate the performance of multiple LLMs, and investigate effective prompt engineering strategies on annotation in historically sensitive contexts. Our multiphase approach combines expert annotation, prompt design, and LLM evaluation using ChatGPT, Llama, and Qwen. We labeled 558 sentences from 15 narrators for sentiment and semantic classification, then developed prompts and evaluated across zero shot, few shot, and retrieval augmented generation (RAG) strategies based on the labeled data. On this benchmark, ChatGPT achieved the highest macro-F1 score for semantic classification (88.71%), followed by Llama (84.99%) and Qwen (83.72%). For sentiment analysis, Llama performed slightly better (82.87%) than Qwen (82.66%) and ChatGPT (82.29%), with all models showing comparable results. Based on the overall evaluation, we selected the best performing prompt configurations for each task and used them to automatically annotate 92,191 sentences from 1,002 interviews in the JAIOH collection. This study develops an LLM-based framework for automatic annotation and analysis of oral history collections. The evaluation results indicate that the proposed LLM-based annotation framework provides promising performance on semantic and sentiment annotation of the JAIOH collection. It contributes a reusable pipeline and practical guidance for applying LLMs in culturally sensitive archival analysis. All code, annotated data, prompt templates, and experimental details are available at the Repository . 1

Authors

Institutions

Publication Details

Journal
Journal on Computing and Cultural Heritage
Published
2026-10-03
DOI
https://doi.org/10.1145/3849086
Primary Topic
Oral History, Memory, Narrative Analysis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Automatic Annotation and Analysis of Oral History using LLMs: An Empirical Study of Densho Digital Collection

Jiangping Chen, Aisa Sakata, Haihua Chen, Komala Subramanyam Cherukuri et al.
Journal on Computing and Cultural Heritage
Oral History, Memory, Narrative Analysis
article

Automatic Annotation and Analysis of Oral History using LLMs: An Empirical Study of Densho Digital Collection

Jiangping Chen, Aisa Sakata, Haihua Chen, Komala Subramanyam Cherukuri, Pranav Abishai Moses
article en

Abstract

Oral histories serve as vital records of lived experience, particularly within communities affected by systemic injustice and historical erasure. Effective and efficient annotation and analysis of oral history archives can promote access and use of the oral histories. However, large scale analysis of these archives remains limited due to their unstructured format, emotional complexity, and the high cost of manual annotation. This paper presents a scalable LLM-based framework to automate semantic and sentiment annotation for oral history archives, with a focus on Japanese American Incarceration Oral History (JAIOH). Using large language models (LLMs), this study seeks to construct a high-quality dataset, systematically evaluate the performance of multiple LLMs, and investigate effective prompt engineering strategies on annotation in historically sensitive contexts. Our multiphase approach combines expert annotation, prompt design, and LLM evaluation using ChatGPT, Llama, and Qwen. We labeled 558 sentences from 15 narrators for sentiment and semantic classification, then developed prompts and evaluated across zero shot, few shot, and retrieval augmented generation (RAG) strategies based on the labeled data. On this benchmark, ChatGPT achieved the highest macro-F1 score for semantic classification (88.71%), followed by Llama (84.99%) and Qwen (83.72%). For sentiment analysis, Llama performed slightly better (82.87%) than Qwen (82.66%) and ChatGPT (82.29%), with all models showing comparable results. Based on the overall evaluation, we selected the best performing prompt configurations for each task and used them to automatically annotate 92,191 sentences from 1,002 interviews in the JAIOH collection. This study develops an LLM-based framework for automatic annotation and analysis of oral history collections. The evaluation results indicate that the proposed LLM-based annotation framework provides promising performance on semantic and sentiment annotation of the JAIOH collection. It contributes a reusable pipeline and practical guidance for applying LLMs in culturally sensitive archival analysis. All code, annotated data, prompt templates, and experimental details are available at the Repository . 1

Journal on Computing and Cultural Heritage
University of North Texas (US), University of Illinois Urbana-Champaign (US)
Openalex Percentile: Top 1%
Oral History, Memory, Narrative Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.