ChatGPT-assisted evaluation of lactational mastitis videos on Douyin: quality and reliability as a health education resource
Short videos have become a primary medium through which the public accesses maternal and child health information, including conditions such as lactational mastitis. With wide reach and visual accessibility, they address the growing demand from patients and their families for evidence-based nursing information. However, the quality of these videos varies considerably, with frequent shortcomings in accuracy, reliability, and overall quality, which may pose a risk of misinformation. Manual assessment tools are difficult to scale to the volume of content on these platforms. Short videos related to lactational mastitis were searched on Douyin (a Chinese short-video platform operated separately from TikTok) from January 2021 to June 2025, and 50 videos were included after strict screening. The artificial intelligence (AI)-DISCERN framework was developed by integrating ChatGPT-4 with the DISCERN instrument and refined through nursing-oriented adaptation and pre-testing (final intraclass correlation coefficient [ICC] = 0.923). AI-assisted scoring was performed using the AI-DISCERN framework. Two trained nursing experts independently performed blinded manual scoring as the reference standard. ICC and weighted Kappa coefficient were used to verify agreement; scoring efficiency was compared, and Spearman rank correlation analysis was used to explore the association between quality and video characteristics. Among the included videos, 86% were posted by professional medical staff and 8% by professional medical institutions, whereas only 3 videos (6%) were posted by non-professional sources. Descriptively, non-professional videos appeared to have higher engagement indicators; however, subgroup findings involving non-professional videos were considered exploratory because of the very small subgroup size. The ICC of the total manual score was 0.976 (95% CI: 0.959–0.986)), indicating excellent agreement; the ICC of the total score between AI and manual scoring was 0.994 (95% CI: 0.989–0.996), and the Kappa value of 12 items was > 0.80. The average time for AI to score a single video was 33.34 ± 7.29 s, which was approximately 2.69 times as fast as manual scoring (89.79 ± 42.09 s, P < 0.001). The total DISCERN score was positively correlated with video duration ( r = 0.498, P < 0.001), but had no significant correlation with the engagement indicators (all P > 0.05). The findings suggest that the AI-DISCERN framework may serve as a feasible supportive tool for the structured assessment of lactational mastitis-related health education materials on social media platforms. Further external validation using larger and more diverse datasets is needed before broader application.
Authors
- B Li
- 时元菊
- Fan Tang (ORCID: https://orcid.org/0000-0002-7218-4959)
- Siyu Zhou
- Xueqin Tang
- Junfeng Li
- Lu An
Institutions
- Guiyang Medical University (CN)
- Affiliated Hospital of Guizhou Medical University (CN)
- Children's Hospital of Chongqing Medical University (CN)
- Chongqing Medical University (CN)
Publication Details
- Journal
- BMC Nursing
- Published
- 2026-09-29
- DOI
- https://doi.org/10.1186/s12912-026-05282-8
- Primary Topic
- Health Literacy and Information Accessibility
- Type
- article
- Field-Weighted Citation Impact
- 0.00