Evaluating the impact of prevalence on the scaled Brier score, Matthews correlation coefficient, and other performance metrics for binary and multiclass hard classifications
Abstract Background Several studies have recently recommended the Matthews correlation coefficient (MCC) to evaluate the accuracy of hard classifications. However, because it is derived from the Pearson correlation, it contains properties that may not be appropriate for classification metrics, such as the value of 0 indicating a lack of relationship. The scaled Brier score was originally conceived for continuous probabilities but may overcome these limitations, such as a value of 0 indicating a random association (i.e., no better than predicting the sample prevalence for each observation). Methods We evaluated the impact of prevalence on the scaled Brier score versus MCC and other common performance metrics with simulated binary and multiclass data. Scenarios with varying sensitivity, specificity, positive predictive value, and prevalence were generated. The following performance metrics were evaluated: accuracy, balanced accuracy, Brier and scaled Brier scores, F score, Kappa and prevalence-adjusted and bias-adjusted Kappa statistics, MCC, receiver operating characteristic and precision-recall area under the curves, and Youden Index. Three prevalence-related criteria were proposed and used to evaluate performance metrics: metric values should not be identical with changes in prevalence; metric values should decrease when prevalence increases and only non-cases are predicted; and metric values should increase when prevalence increases and only cases are predicted. Results Only the scaled Brier score satisfied all three proposed criteria, but accuracy, 1 - Brier score, and PABAK satisfied the first criterion and partially satisfied the second and third criteria. The scaled Brier score consistently reported the lowest values and often deviated substantially from other performance metrics. Conclusion The scaled Brier score was the most consistent with the proposed criteria for binary and multiclass data based on this analysis. The scaled Brier score reported lower values than other performance metrics in all cases and the differences were often large enough to generate substantially different conclusions. The scaled Brier score has previously been recommended as a measure of predictive accuracy for continuous probabilities and the current study indicates that this recommendation can be extended to hard classification data.
Authors
- M Pitz
- Harminder Singh (ORCID: https://orcid.org/0000-0002-9354-2356)
- Pascal Lambert (ORCID: https://orcid.org/0000-0003-2126-7591)
- Kathleen M. Decker
Institutions
- Manitoba Health (CA)
- CancerCare Manitoba (CA)
- Research Institute in Oncology and Hematology (CA)
- University of Manitoba (CA)
Publication Details
- Journal
- Discover Data
- Published
- 2026-09-17
- DOI
- https://doi.org/10.1007/s44248-026-00119-w
- Primary Topic
- Reliability and Agreement in Measurement
- Type
- article
- Field-Weighted Citation Impact
- 0.00