Reliable LLM-Based FinOps Analytics: A Verification Framework for Natural-Language Queries over FOCUS-Compliant Multi-Cloud Billing Data
Multi-cloud environments present a persistent challenge for Financial Operations (FinOps) teams. Billing data from Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP) arrives in proprietary formats with incompatible cost dimensions, and reconciling it takes manual effort and SQL expertise that most financial stakeholders lack. The FinOps Open Cost and Usage Specification (FOCUS) addresses schema fragmentation by defining a vendor-neutral billing data model, and recent work has shown that Large Language Models (LLMs) can translate natural-language billing questions into executable SQL against FOCUS-compliant databases with up to 82.5% accuracy in their original evaluation. For financial decision-making, 82.5% is not enough: a single incorrect aggregation can misstate departmental spend by thousands of dollars. This paper presents the Verification-Augmented FinOps Analytics (VAFA) framework, a multi-stage reliability pipeline that wraps an LLM-based text-to-SQL generator with schema-constrained decoding, round-trip semantic verification, deterministic AST-based validation, and confidence-gated abstention. We evaluate VAFA on a benchmark of 420 natural-language billing questions (120 calibration, 300 held-out test) derived from three cloud provider billing datasets normalized to FOCUS 1.4. The central finding concerns reliability rather than raw text-to-SQL accuracy: through selective abstention, VAFA reduces how often financially consequential errors are accepted without warning. At 81.7% coverage, VAFA achieves 97.1% selective accuracy (95% Wilson CI: 94.2-98.6) with a false-confidence rate of 2.9%, compared with 82.7% selective accuracy at full coverage for the unverified OPTIC baseline. When abstained queries are resolved through clarification, post-clarification accuracy reaches 95.3%. Among abstained queries, 87.3% of clarification requests correctly identified the missing or ambiguous information, so practitioners can correct a query instead of acting on incorrect financial data.
Authors
- Muhammad Zeeshan Hassan
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-10-04
- DOI
- https://doi.org/10.5281/zenodo.23135451
- Primary Topic
- Natural Language Processing Techniques
- Type
- preprint