Regularisation of regression trees by summation of p values

Abstract The standard procedure to decide on the complexity of a CART regression tree is to use cross-validation with the aim of obtaining a predictor that generalises well to unseen data. The randomness in the selection of folds implies that the selected CART regression tree is not a deterministic function of the data. Moreover, the cross-validation procedure may become time consuming and result in inefficient use of training data. We propose a simple deterministic in-sample method that can be used for stopping the growing of a CART regression tree based on node-wise statistical tests. This testing procedure is derived using a connection to change point detection, where the null hypothesis corresponds to no signal. The suggested p value based procedure allows us to consider covariate vectors of arbitrary dimension and allows us to bound the p value of an entire tree from above. Further, we show that the test detects a not too weak signal with a high probability, given a not too small sample size. We illustrate our methodology and the asymptotic results on both simulated and real world data. Additionally, we illustrate how the p value based method can be used to construct a deterministic piece-wise constant auto-calibrated predictor based on a given black-box predictor.

Authors

Institutions

Publication Details

Journal
Computational Statistics
Published
2026-09-19
DOI
https://doi.org/10.1007/s00180-026-01805-8
Primary Topic
Polynomial and algebraic computation
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Regularisation of regression trees by summation of p values

Mathias Lindholm, Nils Engler, Filip Lindskog, Taariq Nazar
Computational Statistics
Polynomial and algebraic computation
article

Regularisation of regression trees by summation of p values

Mathias Lindholm, Nils Engler, Filip Lindskog, Taariq Nazar
article en

Abstract

Abstract The standard procedure to decide on the complexity of a CART regression tree is to use cross-validation with the aim of obtaining a predictor that generalises well to unseen data. The randomness in the selection of folds implies that the selected CART regression tree is not a deterministic function of the data. Moreover, the cross-validation procedure may become time consuming and result in inefficient use of training data. We propose a simple deterministic in-sample method that can be used for stopping the growing of a CART regression tree based on node-wise statistical tests. This testing procedure is derived using a connection to change point detection, where the null hypothesis corresponds to no signal. The suggested p value based procedure allows us to consider covariate vectors of arbitrary dimension and allows us to bound the p value of an entire tree from above. Further, we show that the test detects a not too weak signal with a high probability, given a not too small sample size. We illustrate our methodology and the asymptotic results on both simulated and real world data. Additionally, we illustrate how the p value based method can be used to construct a deterministic piece-wise constant auto-calibrated predictor based on a given black-box predictor.

Computational StatisticsVol. 41(6)
Stockholm University (SE)
Openalex Percentile: Top 97%
Polynomial and algebraic computation
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Regularisation of regression trees by summation of p values — Mathias Lindholm, Nils Engler, et al. · Computational Statistics (2026) | TGRS Research Map | TGRS