An Experimental Evaluation Framework for LLM-Based Multi-Agent Systems in Industrial Contexts: The Kelvin.AI Oil&Gas Case Study
Large language models (LLMs) have enabled significant advancements across industrial operations by automating tasks, providing deeper insights into the data and enhancing productivity. Intelligent assistants provide users with fast answers and valuable information extracted from existing data and documents~\citep{Figlie24}. This paper presents Kelvin.AI Oil\&Gas, a sensor-grounded multi-agent assistant for production engineers for the Oil\&Gas industry and the corresponding evaluation framework. The architecture comprises a main routing agent and eight specialists agents that have access to per-tenant hybrid entity resolution, retrieval-augmented in-context adaptation using automatically generated examples, validated read-only SQL, time-series analytics, structured widgets and interactive clarification. Evaluation combines continuous-integration tests with catalog-derived regression datasets with human curation, and both deterministic and heuristic evaluators. Phoenix is the platform used to conduct all of the tests, where a 131-case baseline achieved 99.2\% response presence, 100\% error-free execution, 88.5\% correct agent selection, and 72.5\% semantic correctness.
Authors
- Ricardo Santos (ORCID: https://orcid.org/0000-0002-2139-5414)
- Fábio Silva (ORCID: https://orcid.org/0000-0001-9872-7117)
- Cláudia Ribeiro
- André Gomes
Institutions
- Universidade do Porto (PT)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-28
- DOI
- https://doi.org/10.5281/zenodo.23022710
- Primary Topic
- Topic Modeling
- Type
- article
- Field-Weighted Citation Impact
- 0.00