Computation of Large Spatial Datasets with the M Function

Since agglomeration is a core question in regional science, spatial concentration measures are widely employed in that field to evaluate the spatial distribution of activities. Increasing access to large geo-referenced datasets, coupled with the development of computing power, has encouraged the search for suitable spatial statistical tools. Distance-based methods have been extensively developed to detect spatial concentration, dispersion or independence of entities at any distance and without any bias. Today, distance-based methods face a new challenge: they must be able to address very large micro-geographic datasets. Recently, Tidu et al. (2024) highlighted the qualities of Marcon and Puech’s M function, a relative distance-based measure, and also expressed reservations about the computation time required. Herein, we explore two possible ways to reduce the computation burden of large geo-located datasets: approximating the position of points and thinning the point pattern. In both cases, the deterioration extent of the M results is estimated and discussed as the gains it provides in computation time, using the R software. We discuss implications of these findings in the field of regional science. We notably provide evidence that the individual location approximation generates information loss at substantially small distances, implying a trade-off between the smallest distance at which spatial interactions could be detected and computing performance. We also give support that random thinning is an efficient method to analyze large datasets with very good accuracy. The R code used in the article is given for the reproducibility of our results.

Authors

Institutions

Publication Details

Journal
International Regional Science Review
Published
2026-09-29
DOI
https://doi.org/10.1177/01600176261492723
Primary Topic
Spatial and Panel Data Analysis
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Computation of Large Spatial Datasets with the M Function

Éric Marcon, Florence Puech
International Regional Science Review
Spatial and Panel Data Analysis
article

Computation of Large Spatial Datasets with the M Function

Éric Marcon, Florence Puech
article en

Abstract

Since agglomeration is a core question in regional science, spatial concentration measures are widely employed in that field to evaluate the spatial distribution of activities. Increasing access to large geo-referenced datasets, coupled with the development of computing power, has encouraged the search for suitable spatial statistical tools. Distance-based methods have been extensively developed to detect spatial concentration, dispersion or independence of entities at any distance and without any bias. Today, distance-based methods face a new challenge: they must be able to address very large micro-geographic datasets. Recently, Tidu et al. (2024) highlighted the qualities of Marcon and Puech’s M function, a relative distance-based measure, and also expressed reservations about the computation time required. Herein, we explore two possible ways to reduce the computation burden of large geo-located datasets: approximating the position of points and thinning the point pattern. In both cases, the deterioration extent of the M results is estimated and discussed as the gains it provides in computation time, using the R software. We discuss implications of these findings in the field of regional science. We notably provide evidence that the individual location approximation generates information loss at substantially small distances, implying a trade-off between the smallest distance at which spatial interactions could be detected and computing performance. We also give support that random thinning is an efficient method to analyze large datasets with very good accuracy. The R code used in the article is given for the reproducibility of our results.

International Regional Science Review
Université de Montpellier (FR), AgroParisTech (FR), Université Paris-Saclay (FR), Institut National de Recherche pour l'Agriculture, l'Alimentation et l'Environnement (FR), Paris-Saclay Applied Economics (FR)
Openalex Percentile: Top 5%
Spatial and Panel Data Analysis
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.