Can an AI Chatbot Reliably Assess the Reporting Quality of Health-Related Qualitative Research? A Pilot Study

Generative AI (genAI) is increasingly being integrated into critical appraisal and knowledge synthesis of research, with growing discourse on its value for qualitative research. Moreover, ChatGPT covers a major portion of the market share and is widely used by researchers and academics. However, a major gap remains in genAI’s ability to assess the reporting quality of qualitative research studies. This pilot study aimed to determine whether GPT-5 could reliably assess the reporting quality of health-related qualitative research studies. The Standards for Reporting Qualitative Research (SRQR) checklist was used as a framework for completing both human and AI chatbot assessments. The SRQR is a widely used validated checklist for enhancing the reporting of qualitative research. Overall, 68 peer-reviewed studies were selected for testing, comparing human SRQR assessments with GPT-5 assessments. Findings revealed a wide variability in agreement between human and GPT-5 assessments, with crude agreement ranging between 60 and 100% and interrater reliability ranging between 0 and 0.85 across checklist items. Simpler, objective checklist items such as study purpose were often correctly identified and had stronger agreement compared to nuanced items. The model struggled with the following SRQR items leading to an “Unclear” response: research design/paradigm (41%), researcher reflexivity (41%), data processing (27%), and trustworthiness (46%). These preliminary results suggest that GPT-5 may struggle with tasks requiring nuance and interpretation of qualitative research, which is rich in complex methods and differing worldviews. Currently, GPT-5 and other AI chatbots should be used with caution and alongside human oversight when assessing the reporting quality of qualitative research within public health and healthcare.

Authors

Institutions

Publication Details

Journal
Knowledge
Published
2026-09-11
DOI
https://doi.org/10.3390/knowledge6030024
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Can an AI Chatbot Reliably Assess the Reporting Quality of Health-Related Qualitative Research? A Pilot Study

Abhinand Thaivalappil, Ian Young, Melissa MacKay, Shan Jin
Knowledge
Artificial Intelligence in Healthcare and Education
article

Can an AI Chatbot Reliably Assess the Reporting Quality of Health-Related Qualitative Research? A Pilot Study

Abhinand Thaivalappil, Ian Young, Melissa MacKay, Shan Jin
article en

Abstract

Generative AI (genAI) is increasingly being integrated into critical appraisal and knowledge synthesis of research, with growing discourse on its value for qualitative research. Moreover, ChatGPT covers a major portion of the market share and is widely used by researchers and academics. However, a major gap remains in genAI’s ability to assess the reporting quality of qualitative research studies. This pilot study aimed to determine whether GPT-5 could reliably assess the reporting quality of health-related qualitative research studies. The Standards for Reporting Qualitative Research (SRQR) checklist was used as a framework for completing both human and AI chatbot assessments. The SRQR is a widely used validated checklist for enhancing the reporting of qualitative research. Overall, 68 peer-reviewed studies were selected for testing, comparing human SRQR assessments with GPT-5 assessments. Findings revealed a wide variability in agreement between human and GPT-5 assessments, with crude agreement ranging between 60 and 100% and interrater reliability ranging between 0 and 0.85 across checklist items. Simpler, objective checklist items such as study purpose were often correctly identified and had stronger agreement compared to nuanced items. The model struggled with the following SRQR items leading to an “Unclear” response: research design/paradigm (41%), researcher reflexivity (41%), data processing (27%), and trustworthiness (46%). These preliminary results suggest that GPT-5 may struggle with tasks requiring nuance and interpretation of qualitative research, which is rich in complex methods and differing worldviews. Currently, GPT-5 and other AI chatbots should be used with caution and alongside human oversight when assessing the reporting quality of qualitative research within public health and healthcare.

KnowledgeVol. 6(3)
Toronto Metropolitan University (CA), University of Guelph (CA)
Openalex Percentile: Top 14%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.