Beyond the gold standard: evaluating AI for subject indexing of Swedish LGBTQ + fiction

Purpose This study evaluates the potential of large language models (LLMs) for subject indexing of Swedish LGBTQ + fiction using the dedicated QLIT thesaurus, based on Homosaurus. It examines both the quantitative performance and qualitative characteristics of AI-generated subject terms, with particular attention to thesaurus compliance and indexing policy. Design/methodology/approach Four general-purpose AI tools (ChatGPT, Claude, Gemini, and DeepSeek) were evaluated on a collection of 14 historical LGBTQ + literary works represented by 16 texts. Subject terms were generated from both metadata records and full texts and compared with existing indexing in the Queerlit database of Swedish LGBTQ + fiction. The metadata subset produced 518 AI-generated subject terms in total, while the full-text subset produced 590 terms. The analysis was based on a coding scheme, which was refined through several rounds of analysis until complete agreement was achieved. Findings Performance was substantially higher for metadata than for full texts, with average F1 scores of 44% and 12%, respectively. Across both datasets, fully incorrect terms constituted the largest category of generated outputs (37.5% for metadata and 56.1% for full texts), while fully correct terms accounted for only 30.0 and 8.3%, respectively. Qualitative analysis revealed recurring problems, including the assignment of broader concepts rather than, or in addition to, more specific concepts, failure to follow thesaurus definitions and indexing policies, confusion between general themes and LGBTQ + -specific themes, and the application of contemporary LGBTQ + concepts to historical texts. Poetry proved particularly challenging because themes were often implicit and open to interpretation. Although a small proportion of generated terms were judged potentially useful despite not appearing in existing indexing, no AI-generated terms were classified as fully correct additions beyond the current Queerlit metadata. Research limitations/implications The study focuses on a relatively specific collection of Swedish historical LGBTQ + fiction and a specialized controlled vocabulary. The findings nevertheless demonstrate the importance of complementing quantitative measures with qualitative evaluation and suggest that assessment of automated subject indexing should consider thesaurus definitions, indexing policies, and interpretive aspects of literary aboutness in addition to agreement with existing metadata. Practical implications The results indicate that current general-purpose LLMs are unlikely to be suitable as semi-automated indexing tools for historical LGBTQ + fiction without substantial human oversight. While AI systems may assist in identifying candidate concepts and additional access points, effective indexing continues to depend on expert knowledge of controlled vocabularies and indexing policy. Originality/value This study contributes to research on AI-assisted subject indexing by combining quantitative and qualitative evaluation of LLM-generated subject terms in a specialized LGBTQ + knowledge organization context. It demonstrates the limitations of gold-standard-based evaluation for fiction indexing and highlights the role of human expertise in applying controlled vocabularies to historically and culturally complex literary materials.

Authors

Institutions

Publication Details

Journal
Journal of Documentation
Published
2026-09-19
DOI
https://doi.org/10.1108/jd-06-2026-0363
Primary Topic
Authorship Attribution and Profiling
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Beyond the gold standard: evaluating AI for subject indexing of Swedish LGBTQ + fiction

Siska Humlesjö, Koraljka Golub, Olof Falk
Journal of Documentation
Authorship Attribution and Profiling
article

Beyond the gold standard: evaluating AI for subject indexing of Swedish LGBTQ + fiction

Siska Humlesjö, Koraljka Golub, Olof Falk
article en

Abstract

Purpose This study evaluates the potential of large language models (LLMs) for subject indexing of Swedish LGBTQ + fiction using the dedicated QLIT thesaurus, based on Homosaurus. It examines both the quantitative performance and qualitative characteristics of AI-generated subject terms, with particular attention to thesaurus compliance and indexing policy. Design/methodology/approach Four general-purpose AI tools (ChatGPT, Claude, Gemini, and DeepSeek) were evaluated on a collection of 14 historical LGBTQ + literary works represented by 16 texts. Subject terms were generated from both metadata records and full texts and compared with existing indexing in the Queerlit database of Swedish LGBTQ + fiction. The metadata subset produced 518 AI-generated subject terms in total, while the full-text subset produced 590 terms. The analysis was based on a coding scheme, which was refined through several rounds of analysis until complete agreement was achieved. Findings Performance was substantially higher for metadata than for full texts, with average F1 scores of 44% and 12%, respectively. Across both datasets, fully incorrect terms constituted the largest category of generated outputs (37.5% for metadata and 56.1% for full texts), while fully correct terms accounted for only 30.0 and 8.3%, respectively. Qualitative analysis revealed recurring problems, including the assignment of broader concepts rather than, or in addition to, more specific concepts, failure to follow thesaurus definitions and indexing policies, confusion between general themes and LGBTQ + -specific themes, and the application of contemporary LGBTQ + concepts to historical texts. Poetry proved particularly challenging because themes were often implicit and open to interpretation. Although a small proportion of generated terms were judged potentially useful despite not appearing in existing indexing, no AI-generated terms were classified as fully correct additions beyond the current Queerlit metadata. Research limitations/implications The study focuses on a relatively specific collection of Swedish historical LGBTQ + fiction and a specialized controlled vocabulary. The findings nevertheless demonstrate the importance of complementing quantitative measures with qualitative evaluation and suggest that assessment of automated subject indexing should consider thesaurus definitions, indexing policies, and interpretive aspects of literary aboutness in addition to agreement with existing metadata. Practical implications The results indicate that current general-purpose LLMs are unlikely to be suitable as semi-automated indexing tools for historical LGBTQ + fiction without substantial human oversight. While AI systems may assist in identifying candidate concepts and additional access points, effective indexing continues to depend on expert knowledge of controlled vocabularies and indexing policy. Originality/value This study contributes to research on AI-assisted subject indexing by combining quantitative and qualitative evaluation of LLM-generated subject terms in a specialized LGBTQ + knowledge organization context. It demonstrates the limitations of gold-standard-based evaluation for fiction indexing and highlights the role of human expertise in applying controlled vocabularies to historically and culturally complex literary materials.

Journal of DocumentationVol. 82(7)
University of Arts and Industrial Design Linz (AT), University of Gothenburg (SE), University of Borås (SE)
Quality Education
Openalex Percentile: Top 8%
Authorship Attribution and Profiling
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.