Exploring new statistical metrics to evaluate the magnitude distribution of earthquake forecasting models

Summary Evaluating earthquake forecasts is a crucial step in understanding and improving the capabilities of forecasting models. The use of specific metrics to assess the consistency between forecasts and data on one particular aspect of the process is important to understand which aspects of seismicity a model is failing to describe, and, consequently, highlight where new versions of the model should improve. This can be done effectively only by metrics unaffected by inconsistencies in other aspects of the process. The Collaboratory for the Study of Earthquake Predictability (CSEP), which organises earthquake forecasting experiments around the globe, has developed different tests targeting different aspects of the process, such as the number N-test or the magnitude M-test, assessing (in)consistency between the observed and forecasted number and magnitude distributions of events, respectively. We find that the results of the recently proposed M-test for catalog-based forecasts (composed of a collection of synthetic catalogs from the model) depend on the N-test, i.e. the two tests do not isolate the desired aspects appropriately, rendering uninformative forecast results. Here, we address this problem using simulated data and provide a solution based on resampling of simulated catalogs. We implement this new M-test with resampling in the pyCSEP software toolkit, conduct two analyses using actual earthquake forecasts for Europe (1990-2015) and Switzerland (1933-1962, 1962-1992, 1993-2022) that provide inconsistent earthquake counts compared to observations, and analyse how the test results change using this proposed test. Lastly, we investigate alternative metrics, namely an unnormalised M-test, two Chi-square formulations, the Hellinger distance, the Brier score, and a novel Multinomial Log-Likelihood (MLL) score, and compare them based on the ability to find (in)consistency between data and forecast in various synthetic scenarios. We find that there are scenarios in which the alternative metrics outperform the resampled M-test. In particular, the MLL test outperforms the M-test in all the scenarios considered. We also study how the ability of finding inconsistency changes with the number of observations, and the cutoff magnitude, comparing one of the Chi-square metrics, the MLL, and the M-test against classical statistical methods such as the Kolmogorov-Smirnoff (KS) test, the Wilcoxon test, and the Anderson-Darling (AD) test. The MLL and the AD test are the ones providing the highest probability of finding inconsistencies for all combinations of number of observations and cutoff magnitude. This study shows how realistic synthetic examples can be used to compare the ability of different metrics in finding inconsistencies between data and forecasts, and shows that the MLL is the metric providing the best tradeoff between interpretability and probability of finding inconsistencies and, therefore, should be used.

Authors

Institutions

Publication Details

Journal
Geophysical Journal International
Published
2026-10-06
DOI
https://doi.org/10.1093/gji/ggag419
Primary Topic
earthquake and tectonic studies
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
article

Exploring new statistical metrics to evaluate the magnitude distribution of earthquake forecasting models

Maximilian Jonas Werner, José A. Bayona, Mark Naylor, Francesco Serafini et al.
Geophysical Journal International
earthquake and tectonic studies
article

Exploring new statistical metrics to evaluate the magnitude distribution of earthquake forecasting models

Maximilian Jonas Werner, José A. Bayona, Mark Naylor, Francesco Serafini, L Mizrahi, K Bayliss, M Han, P Iturrieta
article en

Abstract

Summary Evaluating earthquake forecasts is a crucial step in understanding and improving the capabilities of forecasting models. The use of specific metrics to assess the consistency between forecasts and data on one particular aspect of the process is important to understand which aspects of seismicity a model is failing to describe, and, consequently, highlight where new versions of the model should improve. This can be done effectively only by metrics unaffected by inconsistencies in other aspects of the process. The Collaboratory for the Study of Earthquake Predictability (CSEP), which organises earthquake forecasting experiments around the globe, has developed different tests targeting different aspects of the process, such as the number N-test or the magnitude M-test, assessing (in)consistency between the observed and forecasted number and magnitude distributions of events, respectively. We find that the results of the recently proposed M-test for catalog-based forecasts (composed of a collection of synthetic catalogs from the model) depend on the N-test, i.e. the two tests do not isolate the desired aspects appropriately, rendering uninformative forecast results. Here, we address this problem using simulated data and provide a solution based on resampling of simulated catalogs. We implement this new M-test with resampling in the pyCSEP software toolkit, conduct two analyses using actual earthquake forecasts for Europe (1990-2015) and Switzerland (1933-1962, 1962-1992, 1993-2022) that provide inconsistent earthquake counts compared to observations, and analyse how the test results change using this proposed test. Lastly, we investigate alternative metrics, namely an unnormalised M-test, two Chi-square formulations, the Hellinger distance, the Brier score, and a novel Multinomial Log-Likelihood (MLL) score, and compare them based on the ability to find (in)consistency between data and forecast in various synthetic scenarios. We find that there are scenarios in which the alternative metrics outperform the resampled M-test. In particular, the MLL test outperforms the M-test in all the scenarios considered. We also study how the ability of finding inconsistency changes with the number of observations, and the cutoff magnitude, comparing one of the Chi-square metrics, the MLL, and the M-test against classical statistical methods such as the Kolmogorov-Smirnoff (KS) test, the Wilcoxon test, and the Anderson-Darling (AD) test. The MLL and the AD test are the ones providing the highest probability of finding inconsistencies for all combinations of number of observations and cutoff magnitude. This study shows how realistic synthetic examples can be used to compare the ability of different metrics in finding inconsistencies between data and forecasts, and shows that the MLL is the metric providing the best tradeoff between interpretability and probability of finding inconsistencies and, therefore, should be used.

Geophysical Journal International
ETH Zurich (CH), University of Bristol (GB), Global Earthquake Model (IT), GFZ Helmholtz Centre for Geosciences (DE), University of Edinburgh (GB)
Openalex Percentile: Top 16%
earthquake and tectonic studies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.