Exploring Estimates of Multilevel Reliability for School-Based Behavioral Measures
Many schools utilize universal behavioral screening to quickly evaluate all students to determine who may need additional support. Most commonly, a single teacher rates all students in their class. Like many other school-based assessments, screening data has a nested structure and is typically ordinally scaled and non-normally distributed. In this study, we simulated data with characteristics similar to behavioral screening data with varied numbers of items, levels of inter-item correlation, and average factor loadings. We estimated single-level and multi-level reliability per Lai (Citation2021) to examine the impacts of nesting and ignoring nesting when estimating reliability. With this type of data, single-level alpha and omega were more like within-class reliability estimates than between-class reliability estimates. 5-point ordinal scales were more like the continuous reliability estimates than the 3- or 4-point scales. We discuss the implications of these findings for screening data used in schools and offer recommendations for estimating reliability for future studies of behavioral screening tools.
Authors
- Katie Scarlett Lane Pelton (ORCID: https://orcid.org/0000-0002-7640-6651)
- D. Betsy McCoach (ORCID: https://orcid.org/0000-0001-9063-6835)
Institutions
- Fordham University (US)
- University of Florida (US)
Publication Details
- Journal
- The Journal of Experimental Education
- Published
- 2026-09-15
- DOI
- https://doi.org/10.1080/00220973.2026.2730124
- Primary Topic
- Behavioral and Psychological Studies
- Type
- article
- Field-Weighted Citation Impact
- 0.00