AI versus human coding of NIH grant abstracts

Large language models (LLMs) are increasingly used for qualitative analysis in substance use research, yet their performance relative to human coders remains underexplored. This study compares ChatGPT-4.0 with human coders in performing qualitative coding tasks using NIH grant abstracts as a test case, focusing on the identification and description of research innovations. Using a sample of NIH HEAL Initiative grant abstracts related to opioid overdose prevention, a total of 125 abstracts were independently coded by ChatGPT and humans to generate innovation descriptions, which were then evaluated by both human raters and ChatGPT for depth/detail and relevance/completeness using 5-point Likert scales. Identical instructions were used across all coding and evaluation stages. ChatGPT-generated descriptions were consistently rated higher than human-generated descriptions on both dimensions. Human evaluators rated ChatGPT outputs at an average of 4.47 for both depth/detail and relevance/completeness, compared to 3.33 and 3.24 for human outputs, respectively (F(1,176)=133.9, p < 0.001). These findings suggest that LLMs, when carefully prompted, can enhance the efficiency and quality of qualitative research evaluation.

Authors

Institutions

Publication Details

Journal
PLoS ONE
Published
2026-10-08
DOI
https://doi.org/10.1371/journal.pone.0343857
Primary Topic
Computational and Text Analysis Methods
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

AI versus human coding of NIH grant abstracts

Nadiya Alnoor Jiwa, Todd D. Molfenter, Scott T. Walters, Ginnie Sawyer-Morris et al.
PLoS ONE
Computational and Text Analysis Methods
article

AI versus human coding of NIH grant abstracts

Nadiya Alnoor Jiwa, Todd D. Molfenter, Scott T. Walters, Ginnie Sawyer-Morris, Faye S. Taxman, Justin M. Luningham, Sarah A. Alkhatib, Dallin Judd, Merve Ulukaya
article en

Abstract

Large language models (LLMs) are increasingly used for qualitative analysis in substance use research, yet their performance relative to human coders remains underexplored. This study compares ChatGPT-4.0 with human coders in performing qualitative coding tasks using NIH grant abstracts as a test case, focusing on the identification and description of research innovations. Using a sample of NIH HEAL Initiative grant abstracts related to opioid overdose prevention, a total of 125 abstracts were independently coded by ChatGPT and humans to generate innovation descriptions, which were then evaluated by both human raters and ChatGPT for depth/detail and relevance/completeness using 5-point Likert scales. Identical instructions were used across all coding and evaluation stages. ChatGPT-generated descriptions were consistently rated higher than human-generated descriptions on both dimensions. Human evaluators rated ChatGPT outputs at an average of 4.47 for both depth/detail and relevance/completeness, compared to 3.33 and 3.24 for human outputs, respectively (F(1,176)=133.9, p < 0.001). These findings suggest that LLMs, when carefully prompted, can enhance the efficiency and quality of qualitative research evaluation.

PLoS ONEVol. 21(10)
University of North Texas (US), Texas Christian University (US), University of Wisconsin–Madison (US), George Mason University (US), University of North Texas Health Science Center (US), Friends Research Institute (US)
Openalex Percentile: Top 5%
Computational and Text Analysis Methods
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

AI versus human coding of NIH grant abstracts — Nadiya Alnoor Jiwa, Todd D. Molfenter, et al. · PLoS ONE (2026) | TGRS Research Map | TGRS