Reliable LLM-Based FinOps Analytics: A Verification Framework for Natural-Language Queries over FOCUS-Compliant Multi-Cloud Billing Data

Multi-cloud environments present a persistent challenge for Financial Operations (FinOps) teams. Billing data from Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP) arrives in proprietary formats with incompatible cost dimensions, and reconciling it takes manual effort and SQL expertise that most financial stakeholders lack. The FinOps Open Cost and Usage Specification (FOCUS) addresses schema fragmentation by defining a vendor-neutral billing data model, and recent work has shown that Large Language Models (LLMs) can translate natural-language billing questions into executable SQL against FOCUS-compliant databases with up to 82.5% accuracy in their original evaluation. For financial decision-making, 82.5% is not enough: a single incorrect aggregation can misstate departmental spend by thousands of dollars. This paper presents the Verification-Augmented FinOps Analytics (VAFA) framework, a multi-stage reliability pipeline that wraps an LLM-based text-to-SQL generator with schema-constrained decoding, round-trip semantic verification, deterministic AST-based validation, and confidence-gated abstention. We evaluate VAFA on a benchmark of 420 natural-language billing questions (120 calibration, 300 held-out test) derived from three cloud provider billing datasets normalized to FOCUS 1.4. The central finding concerns reliability rather than raw text-to-SQL accuracy: through selective abstention, VAFA reduces how often financially consequential errors are accepted without warning. At 81.7% coverage, VAFA achieves 97.1% selective accuracy (95% Wilson CI: 94.2-98.6) with a false-confidence rate of 2.9%, compared with 82.7% selective accuracy at full coverage for the unverified OPTIC baseline. When abstained queries are resolved through clarification, post-clarification accuracy reaches 95.3%. Among abstained queries, 87.3% of clarification requests correctly identified the missing or ambiguous information, so practitioners can correct a query instead of acting on incorrect financial data.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-04
DOI
https://doi.org/10.5281/zenodo.23135451
Primary Topic
Natural Language Processing Techniques
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Reliable LLM-Based FinOps Analytics: A Verification Framework for Natural-Language Queries over FOCUS-Compliant Multi-Cloud Billing Data

Muhammad Zeeshan Hassan
Zenodo (CERN European Organization for Nuclear Research)
Natural Language Processing Techniques
preprint

Reliable LLM-Based FinOps Analytics: A Verification Framework for Natural-Language Queries over FOCUS-Compliant Multi-Cloud Billing Data

Muhammad Zeeshan Hassan
preprint en

Abstract

Multi-cloud environments present a persistent challenge for Financial Operations (FinOps) teams. Billing data from Amazon Web Services (AWS), Microsoft Azure, and Google Cloud Platform (GCP) arrives in proprietary formats with incompatible cost dimensions, and reconciling it takes manual effort and SQL expertise that most financial stakeholders lack. The FinOps Open Cost and Usage Specification (FOCUS) addresses schema fragmentation by defining a vendor-neutral billing data model, and recent work has shown that Large Language Models (LLMs) can translate natural-language billing questions into executable SQL against FOCUS-compliant databases with up to 82.5% accuracy in their original evaluation. For financial decision-making, 82.5% is not enough: a single incorrect aggregation can misstate departmental spend by thousands of dollars. This paper presents the Verification-Augmented FinOps Analytics (VAFA) framework, a multi-stage reliability pipeline that wraps an LLM-based text-to-SQL generator with schema-constrained decoding, round-trip semantic verification, deterministic AST-based validation, and confidence-gated abstention. We evaluate VAFA on a benchmark of 420 natural-language billing questions (120 calibration, 300 held-out test) derived from three cloud provider billing datasets normalized to FOCUS 1.4. The central finding concerns reliability rather than raw text-to-SQL accuracy: through selective abstention, VAFA reduces how often financially consequential errors are accepted without warning. At 81.7% coverage, VAFA achieves 97.1% selective accuracy (95% Wilson CI: 94.2-98.6) with a false-confidence rate of 2.9%, compared with 82.7% selective accuracy at full coverage for the unverified OPTIC baseline. When abstained queries are resolved through clarification, post-clarification accuracy reaches 95.3%. Among abstained queries, 87.3% of clarification requests correctly identified the missing or ambiguous information, so practitioners can correct a query instead of acting on incorrect financial data.

Zenodo (CERN European Organization for Nuclear Research)
Natural Language Processing Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Reliable LLM-Based FinOps Analytics: A Verification Framework for Natural-Language Queries over FOCUS-Compliant Multi-Cloud Billing Data — Muhammad Zeeshan Hassan · Zenodo (CERN European Organization for Nuclear Research) (2026) | TGRS Research Map | TGRS