Transfer learning for stochastic frontier models with big data streams
Stochastic frontier analysis separates two-sided noise from one-sided inefficiency, but conventional batch estimation is poorly matched to streams whose frontier and variance composition evolve. We develop FrontierOTL, a prequential online transfer-learning framework that combines a projected stochastic-approximation online expectation-maximization (Online-EM) target expert with frozen source-domain stochastic-frontier experts. Source structure is inferred from a pooled, unlabelled source sample by finite-mixture stochastic-frontier estimation and Bayesian information criterion selection. At each target time, all experts are scored by observable predictive likelihood. A bounded excess negative log-likelihood drives an exponentially weighted aggregation update with a fixed evidence half-life, after which only the target expert is updated. This ordering prevents target leakage and gives an exact log-odds recursion that explains both smoothing and reactivation of previously downweighted experts. We evaluate six estimators in six prespecified data-generating processes that cross stationary, mild gradual and strong abrupt target paths with either transferable source pools or moderately contaminated pools containing 70% related-imperfect and 30% strongly adverse observations. The target expert is strictly cold started: it receives no target observation before the stream begins, and every expert starts with equal weight. Across target sizes 400, 800 and 1,200 and 50 common-random-number replications, FrontierOTL has the lowest mean conditional technical-efficiency root mean squared error, frontier root mean squared error and predictive negative log-likelihood in all 54 scenario-size-metric cells. Among 90 paired comparisons per metric, Holm-adjusted paired t tests are significant in all 90 comparisons for technical efficiency and the frontier and in 72 comparisons for negative log-likelihood; Holm-adjusted Wilcoxon tests are significant in all 90 comparisons for each metric. The BPI Challenge 2019 Purchase-to-Pay application uses a prespecified 26-week historical design period and a subsequent 26-week evaluation stream. FrontierOTL obtains a log-output RMSE of 1.950 and mean predictive NLL of 1.856; local-block-bootstrap LR-type prequential log-score comparisons favour FrontierOTL over all five comparators after Holm correction.
Authors
- Yunquan Song (ORCID: https://orcid.org/0000-0002-4816-8588)
- Xiaorui Bai
Institutions
- China University of Petroleum, East China (CN)
Publication Details
- Journal
- Statistics
- Published
- 2026-09-30
- DOI
- https://doi.org/10.1080/02331888.2026.2735944
- Primary Topic
- Data Stream Mining Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00