Reliability, quality, and readability levels of ChatGPT-5.2 responses regarding office ergonomics: Analysis of content based on Google Trends

BackgroundOffice ergonomics plays a critical role in preventing work-related musculoskeletal disorders and promoting employee well-being. As large language models such as ChatGPT are increasingly used to obtain health-related information, evaluating the quality, reliability, readability, and practical applicability of artificial intelligence (AI)-generated ergonomics content has become increasingly important.ObjectiveTo examine the reliability, quality, accuracy, readability, and clinical applicability levels of the responses given by ChatGPT-5.2 to office ergonomics search terms identified by Google Trends, through expert evaluation.MethodsNineteen keywords were selected from the terms obtained by searching "office ergonomics" on Google Trends on January 7, 2026. Responses were generated individually in incognito mode using ChatGPT-5.2 (December 2025 version) and assessed by two experts (physiotherapist, forensic medicine specialist) using Journal of the American Medical Association Benchmarking Criteria (JAMA), Global Quality Score (GQS), Modified DISCERN Score (MDS), accuracy scale, Flesch-Kincaid Grade Level (FKGL) and Flesch Reading Ease Score (FRES) and Patient Education Material Assessment Tool for printed materials (PEMAT-P); inter-rater agreement was analyzed using Intraclass Correlation Coefficients (ICC) and Cronbach's alpha.ResultsThe mean GQS score was 3.00 ± 0.74; the MDS score was 2.89 ± 0.31; and the accuracy score was 3.58 ± 0.51. The FKGL level was 7.66 ± 2.76 and the FRES score was 54.17 ± 14.91, with 47.4% of the content requiring university-level readability. The PEMAT-P comprehensibility level was 82.53 ± 11.03, and applicability was 49.21 ± 20.76. Inter-rater agreement was good-to-very high (ICC = 0.779-0.974).ConclusionsAlthough ChatGPT-5.2 generally provided acceptable levels of accuracy and content quality in this exploratory analysis, the findings should be interpreted in light of the study's methodological limitations, including the restricted keyword set, single-day assessment, and evaluation of a single model version.

Authors

Institutions

Publication Details

Journal
Work
Published
2026-09-30
DOI
https://doi.org/10.1177/10519815261491758
Primary Topic
Artificial Intelligence in Healthcare and Education
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Reliability, quality, and readability levels of ChatGPT-5.2 responses regarding office ergonomics: Analysis of content based on Google Trends

Ayşe Ünal, Osman Çelbiş
Work
Artificial Intelligence in Healthcare and Education
article

Reliability, quality, and readability levels of ChatGPT-5.2 responses regarding office ergonomics: Analysis of content based on Google Trends

Ayşe Ünal, Osman Çelbiş
article en

Abstract

BackgroundOffice ergonomics plays a critical role in preventing work-related musculoskeletal disorders and promoting employee well-being. As large language models such as ChatGPT are increasingly used to obtain health-related information, evaluating the quality, reliability, readability, and practical applicability of artificial intelligence (AI)-generated ergonomics content has become increasingly important.ObjectiveTo examine the reliability, quality, accuracy, readability, and clinical applicability levels of the responses given by ChatGPT-5.2 to office ergonomics search terms identified by Google Trends, through expert evaluation.MethodsNineteen keywords were selected from the terms obtained by searching "office ergonomics" on Google Trends on January 7, 2026. Responses were generated individually in incognito mode using ChatGPT-5.2 (December 2025 version) and assessed by two experts (physiotherapist, forensic medicine specialist) using Journal of the American Medical Association Benchmarking Criteria (JAMA), Global Quality Score (GQS), Modified DISCERN Score (MDS), accuracy scale, Flesch-Kincaid Grade Level (FKGL) and Flesch Reading Ease Score (FRES) and Patient Education Material Assessment Tool for printed materials (PEMAT-P); inter-rater agreement was analyzed using Intraclass Correlation Coefficients (ICC) and Cronbach's alpha.ResultsThe mean GQS score was 3.00 ± 0.74; the MDS score was 2.89 ± 0.31; and the accuracy score was 3.58 ± 0.51. The FKGL level was 7.66 ± 2.76 and the FRES score was 54.17 ± 14.91, with 47.4% of the content requiring university-level readability. The PEMAT-P comprehensibility level was 82.53 ± 11.03, and applicability was 49.21 ± 20.76. Inter-rater agreement was good-to-very high (ICC = 0.779-0.974).ConclusionsAlthough ChatGPT-5.2 generally provided acceptable levels of accuracy and content quality in this exploratory analysis, the findings should be interpreted in light of the study's methodological limitations, including the restricted keyword set, single-day assessment, and evaluation of a single model version.

Work
Alanya Alaaddin Keykubat Üniversitesi (TR)
Quality Education
Openalex Percentile: Top 16%
Artificial Intelligence in Healthcare and Education
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.