Tri-Modal Calibration for Road Network Conflation: Four-State Labelling, Dirichlet Multi-Class Probabilities, and Conformal Coverage Guarantees
Extended version. This is the full-length (14-page) version of a poster paper accepted to the 34th ACM SIGSPATIAL International Conference on Advances in Geographic Information Systems (SIGSPATIAL '26), Riverside, CA, USA, 3–6 November 2026. A 4-page version appears in the SIGSPATIAL '26 proceedings; this deposit carries the proofs, ablations, and replication results the page limit excluded. Road network conflation matches the links of two road maps, and production pipelines give each match a confidence score that users read as the probability that the match is correct. This paper asks whether that reading is safe. We study 992 hand-labelled matches of TomTom segments to links of KSJ, a Japanese government road network, in three cities of Niigata. Each match is labelled wrong, partial, full or uncertain, where a partial match follows the correct road but overlaps it by at most a fraction τ = 0.60. The class-conditional score distribution is tri-modal: the wrong, partial and full classes peak at different score levels. In the confused band of scores [0.70, 0.85), where auto-accept thresholds typically sit, the matches are 4.7% wrong, 36.8% partial and 58.5% full. For an isotonic calibration baseline fitted on lenient labels, the strict Expected Calibration Error (ECE), which counts partial matches as wrong, is 0.245, and the lenient ECE, which counts them as correct, is 0.019. This gap reveals a calibration failure mode in production pipelines whose per-match scores are read as probabilities. We expose this failure mode with a calibration-free geometric description of partial matches (C0), from which four calibration results follow. (C1) The strict-vs-lenient ECE comparison makes the failure measurable on a single corpus. (C2) Dirichlet calibration reaches a 3-class class-wise ECE of 0.018, 25% below the 0.024 of the uncalibrated baseline. (C3) Isotonic recalibration of the score, fitted to strict labels, reaches a strict ECE of 0.017 ± 0.004 (5-fold cross-validation, 200 seeds). This matches a logistic regression fitted directly on six features with isotonic recalibration (0.014 ± 0.003), so we recommend calibrating the score rather than replacing it. (C4) Split conformal prediction gives a distribution-free coverage guarantee, with a mean empirical coverage of 0.898 ± 0.026 at α = 0.10, against a target of 1 − α = 0.90. Within the literatures we searched over 2024–2026 (scope in the Related Work section), C2 and C4 appear to be the first applications of Dirichlet calibration and split conformal prediction to per-link conflation validity. Data attribution. The evaluated target networks are open-licensed: OpenStreetMap and Overture under ODbL, and the KSJ road network under the MLIT National Land Numerical Information Download Site Content Terms of Use (Government Standard Terms of Use compliant, CC BY 4.0-compatible); use of KSJ data is not endorsed by the Ministry. The TomTom probe feed is used under a commercial licence that permits academic publication of derived results on the licensed three-city polygon but not redistribution of the probe-segment table. v2 (2026-08-26). This version repins every seed-dependent statistic to the 200-seed reporting protocol (previously an unpinned 10-seed set had leaked into several scripts): Table-1 binary ECEs 0.017 ± 0.004 and 0.014 ± 0.003 with Nadeau–Bengio p = 0.791; Hootenanny gaps 4.6× / 3.3×; schema-collapse range [0.008, 0.017]; τ-sweep endpoint 0.293; shape-dispersion AUC lift 0.044 ± 0.024; Venn–Abers 0.021 ± 0.003; Mondrian class-conditional coverage 0.898 ± 0.028. It also corrects the inter-rater audit corpus identity (a separate TomTom-to-KSJ sample, not the open-data corpus), the MCPD sample-count equation, the symmetric shape-dispersion definition, and the CCS concept ids. The 4-page ACM poster (DOI 10.1145/3841645.3843021) reports the corrected values. v3 (2026-10-05). This version revises the wording for readability. Long sentences are split, each technical term is explained where it first appears, and each thing keeps one name: the 992 hand-labelled TomTom-to-KSJ pairs are the flagship corpus, and C1 is called the strict-vs-lenient ECE comparison. Every result of v2 is kept, and the tables are unchanged apart from caption wording. The novelty statements of C2 and C4 now take the hedged, date-bounded form of the 4-page ACM version, and the abstract is shorter. Five statements of v2 are corrected. A sentence that overstated the scope of the KSJ data now says that every result uses KSJ as the target network. References to a 4×4 ablation and a bipartite-assignment baseline that are not in the paper are removed, with the two references cited only there. The evaluation set-up no longer says that all its numbers come from the Niigata polygon. A Python-to-Java ECE difference printed as about 0.03 now reads about 0.017, and the schema-ablation text now gives its table's values (0.017 and 0.008). Some issues are left for a later version: the description of the weighted-sum score as factoring multiplicatively, the normalisation of the overlap fraction (by M rather than M+1 samples), the source given for the lenient ECE of 0.019 in the Hootenanny table, the 10-seed schema ablation described as using the protocol of a 200-seed table, a cross-reference for an optimiser comparison that the cited section does not contain, and a schema-ablation gap printed as 0.227 that rounds to 0.228. The artefact is unchanged from v2.
Authors
- Genaro Peque (ORCID: https://orcid.org/0000-0002-6453-8403)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-05
- DOI
- https://doi.org/10.5281/zenodo.23158777
- Primary Topic
- Automated Road and Building Extraction
- Type
- preprint