A Nonparametric Ordering-Based Regression with Rank-Adaptive Coefficients
This paper introduces Rank-Adaptive Regression, a new estimation framework for continuous outcomes that allows multiple predictor--outcome associations to vary systematically when the dimension organizing such heterogeneity is not known in advance. The method makes the ordering of the population part of the fitting procedure, searching for coefficient combinations that generate predictor-based orderings. This ordering defines a common one-dimensional relative rank dimension, denoted by S∈(0,1], in which lower observed values of the outcome tend to correspond to lower values of S and higher observed values to higher values of S, to the extent attainable from the available predictors. This relationship gives S a direct substantive interpretation, allowing coefficient heterogeneity to be examined across observations that tend to have relatively low, middle, and high outcome values. Estimation combines a stochastic ordering-based search with functional approximation. The first stage estimates Base Coefficients at predefined values of S, while the second uses their means to estimate the Rank-Adaptive Coefficient Functions. Monte Carlo experiments show that the procedure closely approximates the underlying individual ordering: the correlation between the true ranks, Sᵢ, and final estimated ranks, Ŝᵢ, ranges from 0.943 to 0.965 and averages 0.958. The method also reproduces the main constant, linear, and nonlinear patterns of coefficient heterogeneity. An application to the natural logarithm of gross labor income in Bogotá reveals substantial heterogeneity in demographic, educational, household, and employment associations with income across S. For example, an additional year of education is associated with approximately 2.4% higher income at S=0.2 and 33.2% higher income at S=0.8. The Rank-Adaptive specification also improves in-sample fit relative to weighted OLS, increasing the R² from 0.4805 to 0.5188. Rank-Adaptive Regression therefore provides an interpretable framework for characterizing multidimensional coefficient heterogeneity while retaining regression coefficients as the main objects of interpretation.
Authors
- Lina María Sánchez-Céspedes (ORCID: https://orcid.org/0000-0003-0698-8542)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-15
- DOI
- https://doi.org/10.5281/zenodo.22771808
- Primary Topic
- Advanced Statistical Modeling Techniques
- Type
- preprint