HolaGPT Auto: Policy-Level Risk Certification for Cost-Constrained Multi-Model Inference
This working paper proposes a mathematical and operational framework for HolaGPT Auto, a system that selects AI models and execution strategies to minimize total inference cost while meeting explicit requirements for answer reliability, response time, and service coverage. The framework combines capability filtering, contextual outcome prediction, statistical risk certification, and controlled fallback. It evaluates complete execution policies, including verification and answer release, rather than individual model choices alone. The paper includes mathematical derivations, a reproducible synthetic study, and a protocol for future evaluation on the HolaGPT platform. The reported results use artificial data and do not establish performance on real language models or production traffic. Author: Jose Luis Ruedaholagpt.com
Authors
- Jose Luis Rueda
Institutions
- Holistic Management International (US)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-15
- DOI
- https://doi.org/10.5281/zenodo.22761793
- Primary Topic
- Natural Language Processing Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00