A field study on interrater reliability of police-scored Static-99R assessments in sexual assault investigations
Evidence-based policing has increased interest in structured risk assessment to support prevention efforts in sexual and violent crime. This field study examined interrater reliability of police-scored Static-99R assessments in 104 active sexual assault investigations in Alberta, Canada, using police documentation to assess agreement between researcher–researcher, researcher–police, and researcher–civilian administrator rater pairings. Total scores showed excellent consistency across evaluators (ICC = .89–.95), although disagreements sometimes shifted nominal risk categories. Item-level agreement varied widely (κ = –0.09 to 1.00). Exploratory analyses suggested strong researcher–police reliability across two post-training periods following certified instruction. Police total-score agreement with the researcher was comparable for certified versus non-certified instructor training pathways (ICC = .90 versus .88), though subgroup estimates were tempered by sample size. Findings support practical guidance on training models, scorer role assignment, and routine quality assurance (e.g., targeted review of near-threshold classifications) to improve consistency in operational settings.
Authors
- Kevin L. Nunes (ORCID: https://orcid.org/0000-0002-1739-6318)
- Sandy Y. Jung (ORCID: https://orcid.org/0000-0002-9664-3419)
- Heather A. Burke
Institutions
- Carleton University (CA)
- MacEwan University (CA)
Publication Details
- Journal
- Police Practice and Research
- Published
- 2026-10-04
- DOI
- https://doi.org/10.1080/15614263.2026.2741741
- Primary Topic
- Psychopathy, Forensic Psychiatry, Sexual Offending
- Type
- article
- Field-Weighted Citation Impact
- 0.00