Benchmarking Performance Judgment in Open-Weight LLM Controller Tuning: Control Knowledge Does Not Ensure Reliable Gain Ranking

Large language models (LLMs) are increasingly embedded in control-design workflows, yet their ability to compare candidate controllers remains uncertain. A simulator-grounded performance-judgment benchmark evaluates four open-weight checkpoints across a quadruple-tank process, a recycle reactor, and a synthetic 3-by-3 cyclic plant. Declarative control-theory accuracy is nearly saturated (98–100% pooled by plant), whereas qualitative-description accuracy for ranking PI gain sets is 46.8%, 31.8%, and 53.2%, respectively. Outputs are associated with lower gain magnitudes, but the association varies with scaling, plant, and controller sampling. Across 271 replay-stable pairs, IAE rankings agree with ISE rankings on 95.2%; under a combined simulator perturbation, 88.7% of seed–item labels are preserved. On fixed gain pairs, equations and 14-point response trajectories raise pooled accuracy only from 46.8% to 51.2% and 51.0%. All 12 prespecified trajectory-versus-qualitative intervals include zero, and none survives Holm correction. Swapping displayed trajectory values changes 9.5% of parsed choices. By contrast, a numerical scaffold derived from the same samples raises evidence-consistent accuracy from 51.1% to 88.1%, with 11 of 12 contrasts surviving Holm correction. The main bottleneck under these prompts is extracting and aggregating raw numerical evidence, not the final comparison alone. This result does not identify a unique mechanism or establish a general absence of dynamic reasoning. Simulator-based validation remains necessary before LLM judgments are used for controller tuning.

Authors

Institutions

Publication Details

Journal
Processes
Published
2026-09-16
DOI
https://doi.org/10.3390/pr14182954
Primary Topic
Explainable Artificial Intelligence (XAI)
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Benchmarking Performance Judgment in Open-Weight LLM Controller Tuning: Control Knowledge Does Not Ensure Reliable Gain Ranking

Jiaxuan Chen, Haonan Li, Yang Shu
Processes
Explainable Artificial Intelligence (XAI)
article

Benchmarking Performance Judgment in Open-Weight LLM Controller Tuning: Control Knowledge Does Not Ensure Reliable Gain Ranking

Jiaxuan Chen, Haonan Li, Yang Shu
article en

Abstract

Large language models (LLMs) are increasingly embedded in control-design workflows, yet their ability to compare candidate controllers remains uncertain. A simulator-grounded performance-judgment benchmark evaluates four open-weight checkpoints across a quadruple-tank process, a recycle reactor, and a synthetic 3-by-3 cyclic plant. Declarative control-theory accuracy is nearly saturated (98–100% pooled by plant), whereas qualitative-description accuracy for ranking PI gain sets is 46.8%, 31.8%, and 53.2%, respectively. Outputs are associated with lower gain magnitudes, but the association varies with scaling, plant, and controller sampling. Across 271 replay-stable pairs, IAE rankings agree with ISE rankings on 95.2%; under a combined simulator perturbation, 88.7% of seed–item labels are preserved. On fixed gain pairs, equations and 14-point response trajectories raise pooled accuracy only from 46.8% to 51.2% and 51.0%. All 12 prespecified trajectory-versus-qualitative intervals include zero, and none survives Holm correction. Swapping displayed trajectory values changes 9.5% of parsed choices. By contrast, a numerical scaffold derived from the same samples raises evidence-consistent accuracy from 51.1% to 88.1%, with 11 of 12 contrasts surviving Holm correction. The main bottleneck under these prompts is extracting and aggregating raw numerical evidence, not the final comparison alone. This result does not identify a unique mechanism or establish a general absence of dynamic reasoning. Simulator-based validation remains necessary before LLM judgments are used for controller tuning.

ProcessesVol. 14(18)
Jinling Institute of Technology (CN), China Agricultural University (CN), Zhejiang University (CN)
Openalex Percentile: Top 8%
Explainable Artificial Intelligence (XAI)
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.

Benchmarking Performance Judgment in Open-Weight LLM Controller Tuning: Control Knowledge Does Not Ensure Reliable Gain Ranking — Jiaxuan Chen, Haonan Li, et al. · Processes (2026) | TGRS Research Map | TGRS