From detection to action: Using LLM agents for Fault-Tolerant Control

We propose an agentic Large Language Model (LLM) framework for active Fault-Tolerant Control (FTC) that transforms fault detection outputs into constraint-aware recovery actions grounded in plant-specific knowledge. The approach couples (i) a multi-agent workflow that decomposes operator duties into monitoring, planning, action synthesis, simulation, validation, and reprompting; (ii) a Digital Process Plant Twin (DPPT) that exposes plant data, models, and a simulation service for pre-execution testing; and (iii) a Graph Retrieval-Augmented Generation (Graph RAG) layer built on the CPSMod ontology, which organizes plant knowledge (structure, function, hybrid dynamics, control context, and fault semantics) into a graph that supports relation-aware, multi-hop retrieval for the agents. Corrective actions are generated as minimal-risk state-machine recovery paths and corresponding discrete commands or continuous setpoint adaptations, then validated deterministically against interlocks, envelopes, and dynamic feasibility before any actuation. If no acceptable plan is found within a bounded time window, control is handed to a deterministic fallback policy. The framework is evaluated exclusively in simulation on two representative benchmarks: a discrete batch Mixing Module and a Continuous Stirred-Tank Reactor (CSTR) under closed-loop PID regulation. The evaluation isolates the detection-to-action step under a matched-model assumption: the DPPT used for pre-execution validation is identical to the model generating the simulated plant behavior and uses the same injected fault configuration. Robustness to plant-model mismatch and sensor-noise variation or misspecification is not evaluated in this study. Results with lightweight LLMs (GPT-4o-mini and GPT-4.1-mini) show that semantically grounded agents can derive recovery decisions that satisfy the configured validation criteria within latency budgets compatible with the respective process dynamics, demonstrating a pathway from detection to validated corrective action under this matched-model simulation setting.

Authors

Institutions

Publication Details

Journal
Journal of Process Control
Published
2026-09-19
DOI
https://doi.org/10.1016/j.jprocont.2026.103855
Primary Topic
AI-based Problem Solving and Planning
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

From detection to action: Using LLM agents for Fault-Tolerant Control

Artan Markaj, Mehmet Mercangöz, Milapji Singh Gill, Javal Vyas et al.
Journal of Process Control
AI-based Problem Solving and Planning
article

From detection to action: Using LLM agents for Fault-Tolerant Control

Artan Markaj, Mehmet Mercangöz, Milapji Singh Gill, Javal Vyas, Felix Gehlhoff
article en

Abstract

We propose an agentic Large Language Model (LLM) framework for active Fault-Tolerant Control (FTC) that transforms fault detection outputs into constraint-aware recovery actions grounded in plant-specific knowledge. The approach couples (i) a multi-agent workflow that decomposes operator duties into monitoring, planning, action synthesis, simulation, validation, and reprompting; (ii) a Digital Process Plant Twin (DPPT) that exposes plant data, models, and a simulation service for pre-execution testing; and (iii) a Graph Retrieval-Augmented Generation (Graph RAG) layer built on the CPSMod ontology, which organizes plant knowledge (structure, function, hybrid dynamics, control context, and fault semantics) into a graph that supports relation-aware, multi-hop retrieval for the agents. Corrective actions are generated as minimal-risk state-machine recovery paths and corresponding discrete commands or continuous setpoint adaptations, then validated deterministically against interlocks, envelopes, and dynamic feasibility before any actuation. If no acceptable plan is found within a bounded time window, control is handed to a deterministic fallback policy. The framework is evaluated exclusively in simulation on two representative benchmarks: a discrete batch Mixing Module and a Continuous Stirred-Tank Reactor (CSTR) under closed-loop PID regulation. The evaluation isolates the detection-to-action step under a matched-model assumption: the DPPT used for pre-execution validation is identical to the model generating the simulated plant behavior and uses the same injected fault configuration. Robustness to plant-model mismatch and sensor-noise variation or misspecification is not evaluated in this study. Results with lightweight LLMs (GPT-4o-mini and GPT-4.1-mini) show that semantically grounded agents can derive recovery decisions that satisfy the configured validation criteria within latency budgets compatible with the respective process dynamics, demonstrating a pathway from detection to validated corrective action under this matched-model simulation setting.

Journal of Process ControlVol. 167
Helmut Schmidt University (DE), Imperial College London (GB)
Openalex Percentile: Top 41%
AI-based Problem Solving and Planning
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.