Evaluating Intrarater Reliability of Nasopharyngoscopy Ratings Made by Speech-Language Pathologists in a Clinical Versus Research Setting
Purpose: This study aimed to evaluate intrarater reliability for nasopharyngoscopy ratings of velopharyngeal closure across two contexts: ratings made in a clinical setting and ratings made in a research setting using standardized rating definitions. Method: Five speech-language pathologists participated in this study. Two assessments of intrarater reliability were performed: between ratings from clinical reports and re-ratings made in a research setting and videos re-rated twice in a research setting. Nasopharyngoscopy ratings analyzed included velopharyngeal closure pattern, total percent closure, and extent of velar movement. Reliability was calculated using Cohen's kappa and the intraclass correlation coefficient (ICC). Results: Intrarater reliability between clinical reports and re-ratings made in a research setting varied for closure pattern from weak ( k = .24) to perfect ( k = 1.00), for total percent closure from poor (ICC = −.09) to good (ICC = .81), and for extent of velar movement from moderate (ICC = .57) to good (ICC = .80). Compared to intrarater reliability between clinical and research settings, intrarater reliability for videos re-rated twice in a research setting was higher for all raters. For closure pattern, intrarater reliability within a research setting ranged from substantial ( k = .68) to perfect ( k = 1.00), for total percent closure from moderate (ICC = .70) to excellent (ICC = .99), and for extent of velar movement from good (ICC = .80) to excellent (ICC = .91). Conclusions: Findings suggest that nasopharyngoscopy ratings made in a clinical setting may be systematically different than ratings made in a research setting; thus, research results that depend on nasopharyngoscopy ratings may not be generalizable to clinical care. This may be caused by a variety of patient-specific and practice-specific factors, in addition to a lack of standardization. Findings suggest that high reliability for nasopharyngoscopy ratings can be achieved if raters receive consistent training and are following the same rating scale definitions.
Authors
- Jessica L. Chee-Williams (ORCID: https://orcid.org/0000-0002-2882-7275)
- Simone Fischbach
- Jamie L. Perry (ORCID: https://orcid.org/0000-0001-7866-9159)
- Sara Kinter (ORCID: https://orcid.org/0000-0001-6017-4529)
- Thomas J. Sitzman (ORCID: https://orcid.org/0000-0002-2173-923X)
- Katherine Dillon (ORCID: https://orcid.org/0009-0003-0515-8413)
- Adriane L. Baylis (ORCID: https://orcid.org/0000-0003-4148-4729)
- Kelly Nett Cordero (ORCID: https://orcid.org/0000-0001-9426-8638)
- Paula Klaiman
- Megan Donner
Institutions
- Seattle Children's Hospital (US)
- Nationwide Children's Hospital (US)
- University of Arizona (US)
- University of Phoenix (US)
- University of Toronto (CA)
- East Carolina University (US)
- University of Washington (US)
- Hospital for Sick Children (CA)
- Phoenix Children's Hospital (US)
- Mayo Clinic in Arizona (US)
- Children's Healthcare of Atlanta (US)
- The Ohio State University (US)
- Phoenix College (US)
- Arizona State University (US)
- Syracuse University (US)
Publication Details
- Journal
- American Journal of Speech-Language Pathology
- Published
- 2026-09-11
- DOI
- https://doi.org/10.1044/2026_ajslp-26-00234
- Primary Topic
- Cleft Lip and Palate Research
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- National Institute of Dental and Craniofacial Research