Evaluating the impact of prevalence on the scaled Brier score, Matthews correlation coefficient, and other performance metrics for binary and multiclass hard classifications

Abstract Background Several studies have recently recommended the Matthews correlation coefficient (MCC) to evaluate the accuracy of hard classifications. However, because it is derived from the Pearson correlation, it contains properties that may not be appropriate for classification metrics, such as the value of 0 indicating a lack of relationship. The scaled Brier score was originally conceived for continuous probabilities but may overcome these limitations, such as a value of 0 indicating a random association (i.e., no better than predicting the sample prevalence for each observation). Methods We evaluated the impact of prevalence on the scaled Brier score versus MCC and other common performance metrics with simulated binary and multiclass data. Scenarios with varying sensitivity, specificity, positive predictive value, and prevalence were generated. The following performance metrics were evaluated: accuracy, balanced accuracy, Brier and scaled Brier scores, F score, Kappa and prevalence-adjusted and bias-adjusted Kappa statistics, MCC, receiver operating characteristic and precision-recall area under the curves, and Youden Index. Three prevalence-related criteria were proposed and used to evaluate performance metrics: metric values should not be identical with changes in prevalence; metric values should decrease when prevalence increases and only non-cases are predicted; and metric values should increase when prevalence increases and only cases are predicted. Results Only the scaled Brier score satisfied all three proposed criteria, but accuracy, 1 - Brier score, and PABAK satisfied the first criterion and partially satisfied the second and third criteria. The scaled Brier score consistently reported the lowest values and often deviated substantially from other performance metrics. Conclusion The scaled Brier score was the most consistent with the proposed criteria for binary and multiclass data based on this analysis. The scaled Brier score reported lower values than other performance metrics in all cases and the differences were often large enough to generate substantially different conclusions. The scaled Brier score has previously been recommended as a measure of predictive accuracy for continuous probabilities and the current study indicates that this recommendation can be extended to hard classification data.

Authors

Institutions

Publication Details

Journal
Discover Data
Published
2026-09-17
DOI
https://doi.org/10.1007/s44248-026-00119-w
Primary Topic
Reliability and Agreement in Measurement
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Evaluating the impact of prevalence on the scaled Brier score, Matthews correlation coefficient, and other performance metrics for binary and multiclass hard classifications

M Pitz, Harminder Singh, Pascal Lambert, Kathleen M. Decker
Discover Data
Reliability and Agreement in Measurement
article

Evaluating the impact of prevalence on the scaled Brier score, Matthews correlation coefficient, and other performance metrics for binary and multiclass hard classifications

M Pitz, Harminder Singh, Pascal Lambert, Kathleen M. Decker
article en

Abstract

Abstract Background Several studies have recently recommended the Matthews correlation coefficient (MCC) to evaluate the accuracy of hard classifications. However, because it is derived from the Pearson correlation, it contains properties that may not be appropriate for classification metrics, such as the value of 0 indicating a lack of relationship. The scaled Brier score was originally conceived for continuous probabilities but may overcome these limitations, such as a value of 0 indicating a random association (i.e., no better than predicting the sample prevalence for each observation). Methods We evaluated the impact of prevalence on the scaled Brier score versus MCC and other common performance metrics with simulated binary and multiclass data. Scenarios with varying sensitivity, specificity, positive predictive value, and prevalence were generated. The following performance metrics were evaluated: accuracy, balanced accuracy, Brier and scaled Brier scores, F score, Kappa and prevalence-adjusted and bias-adjusted Kappa statistics, MCC, receiver operating characteristic and precision-recall area under the curves, and Youden Index. Three prevalence-related criteria were proposed and used to evaluate performance metrics: metric values should not be identical with changes in prevalence; metric values should decrease when prevalence increases and only non-cases are predicted; and metric values should increase when prevalence increases and only cases are predicted. Results Only the scaled Brier score satisfied all three proposed criteria, but accuracy, 1 - Brier score, and PABAK satisfied the first criterion and partially satisfied the second and third criteria. The scaled Brier score consistently reported the lowest values and often deviated substantially from other performance metrics. Conclusion The scaled Brier score was the most consistent with the proposed criteria for binary and multiclass data based on this analysis. The scaled Brier score reported lower values than other performance metrics in all cases and the differences were often large enough to generate substantially different conclusions. The scaled Brier score has previously been recommended as a measure of predictive accuracy for continuous probabilities and the current study indicates that this recommendation can be extended to hard classification data.

Discover DataVol. 4(1)
Manitoba Health (CA), CancerCare Manitoba (CA), Research Institute in Oncology and Hematology (CA), University of Manitoba (CA)
Openalex Percentile: Top 9%
Reliability and Agreement in Measurement
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.