Transfer learning for stochastic frontier models with big data streams

Stochastic frontier analysis separates two-sided noise from one-sided inefficiency, but conventional batch estimation is poorly matched to streams whose frontier and variance composition evolve. We develop FrontierOTL, a prequential online transfer-learning framework that combines a projected stochastic-approximation online expectation-maximization (Online-EM) target expert with frozen source-domain stochastic-frontier experts. Source structure is inferred from a pooled, unlabelled source sample by finite-mixture stochastic-frontier estimation and Bayesian information criterion selection. At each target time, all experts are scored by observable predictive likelihood. A bounded excess negative log-likelihood drives an exponentially weighted aggregation update with a fixed evidence half-life, after which only the target expert is updated. This ordering prevents target leakage and gives an exact log-odds recursion that explains both smoothing and reactivation of previously downweighted experts. We evaluate six estimators in six prespecified data-generating processes that cross stationary, mild gradual and strong abrupt target paths with either transferable source pools or moderately contaminated pools containing 70% related-imperfect and 30% strongly adverse observations. The target expert is strictly cold started: it receives no target observation before the stream begins, and every expert starts with equal weight. Across target sizes 400, 800 and 1,200 and 50 common-random-number replications, FrontierOTL has the lowest mean conditional technical-efficiency root mean squared error, frontier root mean squared error and predictive negative log-likelihood in all 54 scenario-size-metric cells. Among 90 paired comparisons per metric, Holm-adjusted paired t tests are significant in all 90 comparisons for technical efficiency and the frontier and in 72 comparisons for negative log-likelihood; Holm-adjusted Wilcoxon tests are significant in all 90 comparisons for each metric. The BPI Challenge 2019 Purchase-to-Pay application uses a prespecified 26-week historical design period and a subsequent 26-week evaluation stream. FrontierOTL obtains a log-output RMSE of 1.950 and mean predictive NLL of 1.856; local-block-bootstrap LR-type prequential log-score comparisons favour FrontierOTL over all five comparators after Holm correction.

Authors

Institutions

Publication Details

Journal
Statistics
Published
2026-09-30
DOI
https://doi.org/10.1080/02331888.2026.2735944
Primary Topic
Data Stream Mining Techniques
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Transfer learning for stochastic frontier models with big data streams

Yunquan Song, Xiaorui Bai
Statistics
Data Stream Mining Techniques
article

Transfer learning for stochastic frontier models with big data streams

Yunquan Song, Xiaorui Bai
article en

Abstract

Stochastic frontier analysis separates two-sided noise from one-sided inefficiency, but conventional batch estimation is poorly matched to streams whose frontier and variance composition evolve. We develop FrontierOTL, a prequential online transfer-learning framework that combines a projected stochastic-approximation online expectation-maximization (Online-EM) target expert with frozen source-domain stochastic-frontier experts. Source structure is inferred from a pooled, unlabelled source sample by finite-mixture stochastic-frontier estimation and Bayesian information criterion selection. At each target time, all experts are scored by observable predictive likelihood. A bounded excess negative log-likelihood drives an exponentially weighted aggregation update with a fixed evidence half-life, after which only the target expert is updated. This ordering prevents target leakage and gives an exact log-odds recursion that explains both smoothing and reactivation of previously downweighted experts. We evaluate six estimators in six prespecified data-generating processes that cross stationary, mild gradual and strong abrupt target paths with either transferable source pools or moderately contaminated pools containing 70% related-imperfect and 30% strongly adverse observations. The target expert is strictly cold started: it receives no target observation before the stream begins, and every expert starts with equal weight. Across target sizes 400, 800 and 1,200 and 50 common-random-number replications, FrontierOTL has the lowest mean conditional technical-efficiency root mean squared error, frontier root mean squared error and predictive negative log-likelihood in all 54 scenario-size-metric cells. Among 90 paired comparisons per metric, Holm-adjusted paired t tests are significant in all 90 comparisons for technical efficiency and the frontier and in 72 comparisons for negative log-likelihood; Holm-adjusted Wilcoxon tests are significant in all 90 comparisons for each metric. The BPI Challenge 2019 Purchase-to-Pay application uses a prespecified 26-week historical design period and a subsequent 26-week evaluation stream. FrontierOTL obtains a log-output RMSE of 1.950 and mean predictive NLL of 1.856; local-block-bootstrap LR-type prequential log-score comparisons favour FrontierOTL over all five comparators after Holm correction.

Statistics
China University of Petroleum, East China (CN)
Openalex Percentile: Top 9%
Data Stream Mining Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.