Keeping the Whole Model Family: Bootstrap-Aggregated Phenomenological Ensembles for Mining and Industrial Process Data

Industrial practice fits a family of candidate phenomenological models to process data and then keeps one winner, selected by R2 or an information criterion. That discards the evidence about model form and leaves the reported prediction with no honest uncertainty. The flotation literature itself has documented the consequence: fitted rate-constant distributions are often artifacts of a two-parameter constraint, and the estimated maximum recovery and rate distribution of a batch test move substantially when a single sampling time is removed. This preprint defines BAPE, bootstrap-aggregated phenomenological ensembles: the random-forest recipe with curated closed-form process models as base learners. Each ensemble member sees a bootstrap resample of the data AND a random subset of the model library, fits every family in that subset, and keeps the information-criterion best; the population of members then yields three distinct readings, a calibrated predictive distribution, per-family inclusion probabilities, and the parameter clouds that make equifinality visible. The library-subsampling axis is the direct analogue of feature subsampling in a random forest, applied to equation structure rather than covariates. A bank of 42 cited families is implemented across flotation (batch and continuous), comminution energy laws and batch population balances, thickening, leaching, thermal derating and plant utilities, and twelve methods are compared on the same held-out protocol: single best fit, information-criterion selection and averaging, GLUE, single-structure bootstrap bagging, BAPE, cross-validated stacking, Bayesian model averaging, Ensemble-SINDy as the generic-library contrast, a Kennedy-O'Hagan Gaussian-process discrepancy hybrid, deep ensembles as the black-box control, and a mixture of phenomenological experts whose gate weights frozen closed-form curves. The case matrix spans real and synthetic data: a real 737,453-row iron ore flotation circuit reduced through a documented leakage gate, nineteen digitized published settling series, eleven digitized batch flotation series, eleven years of measured national mining water use, the national mining electricity record, two industrial transfer records, mechanism-truth cases generated by a sedimentation partial differential equation, a radial leaching column and a multi-size population balance, and two designed controls. Results are reported as unit-free rank aggregations because the observables span recovery fractions, megawatts, terawatt-hours and litres per second. Over 1,329 method-variant rows, six ensemble rungs place ahead of the single-best-fit control under extrapolation; Bayesian model averaging is the best calibrated (interval-score rank 3.46, coverage 0.83 against a nominal 0.90); a mixture of phenomenological experts identifies the generating family most often (55.6 percent exact recovery); and BAPE itself lands level with the control on point accuracy while ranking second on calibration, a result that moved across three successive matrices and is reported rather than suppressed. Broken down by unit process, no single rung wins everywhere: gains over the control range from 86 percent on a steel emission-factor record to zero on the real flotation plant, and BAPE leads the nineteen published settling series by 16 percent. The sharpest industrial result comes from the real circuit: its residence time varies by about five percent while recovery moves ten points, so flotation kinetics are not identifiable from routine operating data and equifinality, not a rate constant, is the correct answer. Engine (MIT, PyPI): https://pypi.org/project/phenoforge/ ; https://github.com/fsantibanezleal/CAOS_PhenoForge . Live instance: https://fragua.ml.fasl-work.com .

Authors

Institutions

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-08-28
DOI
https://doi.org/10.5281/zenodo.22144371
Primary Topic
Mineral Processing and Grinding
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
preprint

Keeping the Whole Model Family: Bootstrap-Aggregated Phenomenological Ensembles for Mining and Industrial Process Data

Felipe Santibañez-Leal
Zenodo (CERN European Organization for Nuclear Research)
Mineral Processing and Grinding
preprint

Keeping the Whole Model Family: Bootstrap-Aggregated Phenomenological Ensembles for Mining and Industrial Process Data

Felipe Santibañez-Leal
preprint en

Abstract

Industrial practice fits a family of candidate phenomenological models to process data and then keeps one winner, selected by R2 or an information criterion. That discards the evidence about model form and leaves the reported prediction with no honest uncertainty. The flotation literature itself has documented the consequence: fitted rate-constant distributions are often artifacts of a two-parameter constraint, and the estimated maximum recovery and rate distribution of a batch test move substantially when a single sampling time is removed. This preprint defines BAPE, bootstrap-aggregated phenomenological ensembles: the random-forest recipe with curated closed-form process models as base learners. Each ensemble member sees a bootstrap resample of the data AND a random subset of the model library, fits every family in that subset, and keeps the information-criterion best; the population of members then yields three distinct readings, a calibrated predictive distribution, per-family inclusion probabilities, and the parameter clouds that make equifinality visible. The library-subsampling axis is the direct analogue of feature subsampling in a random forest, applied to equation structure rather than covariates. A bank of 42 cited families is implemented across flotation (batch and continuous), comminution energy laws and batch population balances, thickening, leaching, thermal derating and plant utilities, and twelve methods are compared on the same held-out protocol: single best fit, information-criterion selection and averaging, GLUE, single-structure bootstrap bagging, BAPE, cross-validated stacking, Bayesian model averaging, Ensemble-SINDy as the generic-library contrast, a Kennedy-O'Hagan Gaussian-process discrepancy hybrid, deep ensembles as the black-box control, and a mixture of phenomenological experts whose gate weights frozen closed-form curves. The case matrix spans real and synthetic data: a real 737,453-row iron ore flotation circuit reduced through a documented leakage gate, nineteen digitized published settling series, eleven digitized batch flotation series, eleven years of measured national mining water use, the national mining electricity record, two industrial transfer records, mechanism-truth cases generated by a sedimentation partial differential equation, a radial leaching column and a multi-size population balance, and two designed controls. Results are reported as unit-free rank aggregations because the observables span recovery fractions, megawatts, terawatt-hours and litres per second. Over 1,329 method-variant rows, six ensemble rungs place ahead of the single-best-fit control under extrapolation; Bayesian model averaging is the best calibrated (interval-score rank 3.46, coverage 0.83 against a nominal 0.90); a mixture of phenomenological experts identifies the generating family most often (55.6 percent exact recovery); and BAPE itself lands level with the control on point accuracy while ranking second on calibration, a result that moved across three successive matrices and is reported rather than suppressed. Broken down by unit process, no single rung wins everywhere: gains over the control range from 86 percent on a steel emission-factor record to zero on the real flotation plant, and BAPE leads the nineteen published settling series by 16 percent. The sharpest industrial result comes from the real circuit: its residence time varies by about five percent while recovery moves ten points, so flotation kinetics are not identifiable from routine operating data and equifinality, not a rate constant, is the correct answer. Engine (MIT, PyPI): https://pypi.org/project/phenoforge/ ; https://github.com/fsantibanezleal/CAOS_PhenoForge . Live instance: https://fragua.ml.fasl-work.com .

Zenodo (CERN European Organization for Nuclear Research)
Open University of Cyprus (CY)
Mineral Processing and Grinding
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.