Evaluating Intrarater Reliability of Nasopharyngoscopy Ratings Made by Speech-Language Pathologists in a Clinical Versus Research Setting

Purpose: This study aimed to evaluate intrarater reliability for nasopharyngoscopy ratings of velopharyngeal closure across two contexts: ratings made in a clinical setting and ratings made in a research setting using standardized rating definitions. Method: Five speech-language pathologists participated in this study. Two assessments of intrarater reliability were performed: between ratings from clinical reports and re-ratings made in a research setting and videos re-rated twice in a research setting. Nasopharyngoscopy ratings analyzed included velopharyngeal closure pattern, total percent closure, and extent of velar movement. Reliability was calculated using Cohen's kappa and the intraclass correlation coefficient (ICC). Results: Intrarater reliability between clinical reports and re-ratings made in a research setting varied for closure pattern from weak ( k = .24) to perfect ( k = 1.00), for total percent closure from poor (ICC = −.09) to good (ICC = .81), and for extent of velar movement from moderate (ICC = .57) to good (ICC = .80). Compared to intrarater reliability between clinical and research settings, intrarater reliability for videos re-rated twice in a research setting was higher for all raters. For closure pattern, intrarater reliability within a research setting ranged from substantial ( k = .68) to perfect ( k = 1.00), for total percent closure from moderate (ICC = .70) to excellent (ICC = .99), and for extent of velar movement from good (ICC = .80) to excellent (ICC = .91). Conclusions: Findings suggest that nasopharyngoscopy ratings made in a clinical setting may be systematically different than ratings made in a research setting; thus, research results that depend on nasopharyngoscopy ratings may not be generalizable to clinical care. This may be caused by a variety of patient-specific and practice-specific factors, in addition to a lack of standardization. Findings suggest that high reliability for nasopharyngoscopy ratings can be achieved if raters receive consistent training and are following the same rating scale definitions.

Authors

Institutions

Publication Details

Journal
American Journal of Speech-Language Pathology
Published
2026-09-11
DOI
https://doi.org/10.1044/2026_ajslp-26-00234
Primary Topic
Cleft Lip and Palate Research
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evaluating Intrarater Reliability of Nasopharyngoscopy Ratings Made by Speech-Language Pathologists in a Clinical Versus Research Setting

Jessica L. Chee-Williams, Simone Fischbach, Jamie L. Perry, Sara Kinter et al.
American Journal of Speech-Language Pathology
Cleft Lip and Palate Research
article

Evaluating Intrarater Reliability of Nasopharyngoscopy Ratings Made by Speech-Language Pathologists in a Clinical Versus Research Setting

Jessica L. Chee-Williams, Simone Fischbach, Jamie L. Perry, Sara Kinter, Thomas J. Sitzman, Katherine Dillon, Adriane L. Baylis, Kelly Nett Cordero, Paula Klaiman, Megan Donner
article en

Abstract

Purpose: This study aimed to evaluate intrarater reliability for nasopharyngoscopy ratings of velopharyngeal closure across two contexts: ratings made in a clinical setting and ratings made in a research setting using standardized rating definitions. Method: Five speech-language pathologists participated in this study. Two assessments of intrarater reliability were performed: between ratings from clinical reports and re-ratings made in a research setting and videos re-rated twice in a research setting. Nasopharyngoscopy ratings analyzed included velopharyngeal closure pattern, total percent closure, and extent of velar movement. Reliability was calculated using Cohen's kappa and the intraclass correlation coefficient (ICC). Results: Intrarater reliability between clinical reports and re-ratings made in a research setting varied for closure pattern from weak ( k = .24) to perfect ( k = 1.00), for total percent closure from poor (ICC = −.09) to good (ICC = .81), and for extent of velar movement from moderate (ICC = .57) to good (ICC = .80). Compared to intrarater reliability between clinical and research settings, intrarater reliability for videos re-rated twice in a research setting was higher for all raters. For closure pattern, intrarater reliability within a research setting ranged from substantial ( k = .68) to perfect ( k = 1.00), for total percent closure from moderate (ICC = .70) to excellent (ICC = .99), and for extent of velar movement from good (ICC = .80) to excellent (ICC = .91). Conclusions: Findings suggest that nasopharyngoscopy ratings made in a clinical setting may be systematically different than ratings made in a research setting; thus, research results that depend on nasopharyngoscopy ratings may not be generalizable to clinical care. This may be caused by a variety of patient-specific and practice-specific factors, in addition to a lack of standardization. Findings suggest that high reliability for nasopharyngoscopy ratings can be achieved if raters receive consistent training and are following the same rating scale definitions.

American Journal of Speech-Language Pathology
Seattle Children's Hospital (US), Nationwide Children's Hospital (US), University of Arizona (US), University of Phoenix (US), University of Toronto (CA), East Carolina University (US), University of Washington (US), Hospital for Sick Children (CA), Phoenix Children's Hospital (US), Mayo Clinic in Arizona (US), Children's Healthcare of Atlanta (US), The Ohio State University (US), Phoenix College (US), Arizona State University (US), Syracuse University (US)
National Institute of Dental and Craniofacial Research
No poverty
Openalex Percentile: Top 11%
Cleft Lip and Palate Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.