Development, preliminary measurement evaluation, and longitudinal application of a competency assessment framework for newly recruited nurse anesthetists
Evidence remains limited on how specialty-specific competency scores change month by month during the transition of newly recruited nurses into anesthesia practice and on whether self- and faculty-rating trajectories differ. We therefore developed a role-specific framework and examined preliminary content-related evidence, total-score inter-rater reliability, concurrent associations, exploratory T0–T6 change, and longitudinal score patterns. This single-centre, multi-stage study integrated literature review and expert interviews, two-round Delphi consultation, analytic hierarchy process weighting, preliminary measurement evaluation, and retrospective analysis of a census of eligible nurses entering the programme from August 2016 to August 2025. T0 denoted programme entry after pre-employment preparation; T1–T6 denoted completion of months 1–6. The longitudinal analysis used complete T0–T6 records. Repeated-measures ANOVA was the primary within-evaluator time analysis, and participant-level mixed-effects models compared evaluator trajectories. Analyses were performed in R 4.5.2. Of 77 screened records, 68 (88.3%) had complete T0–T6 data and were analysed; five nurses resigned and four had at least two consecutive missing monthly assessments. The framework comprised five domains, 20 s-level indicators, and 100 scored elements. S-CVI/Ave was 0.95 and item-level modified kappa ranged from 0.831 to 1.000, reflecting expert relevance ratings and chance-corrected agreement, respectively. In five nurses rated independently by two prespecified evaluators, total-score ICC(A,1) at T0 was 0.889 (95% CI 0.388–0.988). Mean self-assessment scores increased from 35.95 at T0 to 96.83 at T6 and faculty scores from 29.20 to 95.32; Greenhouse–Geisser-corrected time effects were observed for both perspectives (both P < 0.001; partial η²=0.977 and 0.984). Self-ratings exceeded faculty ratings at every time point; the model-based difference was largest at T2 (9.51 points) and narrowed to 1.50 points at T6 (time-by-evaluator LR χ²=110.67, df = 6, P < 0.001). At T3, self- and faculty scores correlated with written examination scores at ρ = 0.815 and 0.582 and with OSCE scores at ρ = 0.580 and 0.558, respectively; at T6, the corresponding correlations were 0.752 and 0.532 for written examination and 0.689 and 0.340 for OSCE. T0 written-examination correlations were negligible. In a separate prespecified five-person subset used for an exploratory T0–T6 change analysis, all T0–T6 changes were positive, but the exact two-sided sign-test P value was 0.0625. Scores increased longitudinally, self-ratings remained higher than faculty ratings but converged over time, and both rating sources showed positive concurrent associations with written examination and OSCE scores at T3 and T6. These results provide preliminary, low-precision measurement evidence rather than comprehensive validation or causal evidence of training effects; prospective multicentre evaluation is required before high-stakes use.
Authors
- Xiaobei Ma
- 高凤莉
- Xin Liu
- Zhifeng Gao
- Yi Duan
- Xueyan Fan
Institutions
- Beijing Tsinghua Chang Gung Hospital (CN)
- Tsinghua University (CN)
Publication Details
- Journal
- BMC Nursing
- Published
- 2026-09-28
- DOI
- https://doi.org/10.1186/s12912-026-05424-y
- Primary Topic
- Innovations in Medical Education
- Type
- article
- Field-Weighted Citation Impact
- 0.00