Keeping the Whole Model Family: Bootstrap-Aggregated Phenomenological Ensembles for Mining and Industrial Process Data
Industrial practice fits a family of candidate phenomenological models to process data and then keeps one winner, selected by R2 or an information criterion. That discards the evidence about model form and leaves the reported prediction with no honest uncertainty. The flotation literature itself has documented the consequence: fitted rate-constant distributions are often artifacts of a two-parameter constraint, and the estimated maximum recovery and rate distribution of a batch test move substantially when a single sampling time is removed. This preprint defines BAPE, bootstrap-aggregated phenomenological ensembles: the random-forest recipe with curated closed-form process models as base learners. Each ensemble member sees a bootstrap resample of the data AND a random subset of the model library, fits every family in that subset, and keeps the information-criterion best; the population of members then yields three distinct readings, a calibrated predictive distribution, per-family inclusion probabilities, and the parameter clouds that make equifinality visible. The library-subsampling axis is the direct analogue of feature subsampling in a random forest, applied to equation structure rather than covariates. A bank of 42 cited families is implemented across flotation (batch and continuous), comminution energy laws and batch population balances, thickening, leaching, thermal derating and plant utilities, and twelve methods are compared on the same held-out protocol: single best fit, information-criterion selection and averaging, GLUE, single-structure bootstrap bagging, BAPE, cross-validated stacking, Bayesian model averaging, Ensemble-SINDy as the generic-library contrast, a Kennedy-O'Hagan Gaussian-process discrepancy hybrid, deep ensembles as the black-box control, and a mixture of phenomenological experts whose gate weights frozen closed-form curves. The case matrix spans real and synthetic data: a real 737,453-row iron ore flotation circuit reduced through a documented leakage gate, nineteen digitized published settling series, eleven digitized batch flotation series, eleven years of measured national mining water use, the national mining electricity record, two industrial transfer records, mechanism-truth cases generated by a sedimentation partial differential equation, a radial leaching column and a multi-size population balance, and two designed controls. Results are reported as unit-free rank aggregations because the observables span recovery fractions, megawatts, terawatt-hours and litres per second. Over 1,329 method-variant rows, six ensemble rungs place ahead of the single-best-fit control under extrapolation; Bayesian model averaging is the best calibrated (interval-score rank 3.46, coverage 0.83 against a nominal 0.90); a mixture of phenomenological experts identifies the generating family most often (55.6 percent exact recovery); and BAPE itself lands level with the control on point accuracy while ranking second on calibration, a result that moved across three successive matrices and is reported rather than suppressed. Broken down by unit process, no single rung wins everywhere: gains over the control range from 86 percent on a steel emission-factor record to zero on the real flotation plant, and BAPE leads the nineteen published settling series by 16 percent. The sharpest industrial result comes from the real circuit: its residence time varies by about five percent while recovery moves ten points, so flotation kinetics are not identifiable from routine operating data and equifinality, not a rate constant, is the correct answer. Engine (MIT, PyPI): https://pypi.org/project/phenoforge/ ; https://github.com/fsantibanezleal/CAOS_PhenoForge . Live instance: https://fragua.ml.fasl-work.com .
Authors
- Felipe Santibañez-Leal (ORCID: https://orcid.org/0000-0002-0150-3246)
Institutions
- Open University of Cyprus (CY)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-08-28
- DOI
- https://doi.org/10.5281/zenodo.22144371
- Primary Topic
- Mineral Processing and Grinding
- Type
- preprint