ReliHarness: A Self-Learning Reliability Harness for LLM Agent Tool Execution

Large language model agents invoke external tools through the Model Context Protocol to perform actions beyond text generation. Once a background thread is dispatched, even the latest Model Context Protocol Tasks extension exposes only coarse-grained, self-reported task states, so a thread that deadlocks or crashes silently remains reported as working and invisible to the client, which is not acceptable in some industrial applications that need long monitoring. We present ReliHarness, a self-learning dual-channel execution-layer reliability harness framework that instead probes each worker thread at a lower level through two independent checks: heartbeat timeout detection and operating system-level thread liveness detection. The proposed framework contributes a five-state lifecycle machine, dual-channel detection, bi-layer evidence logging, a four-stage cascade pipeline, and a self-learning command-matching module. Experiments on a Phytium S5000C and Ascend 310P platform achieve 100% detection coverage for the tested crash and deadlock scenarios with zero false positives. The self-learning module expands the knowledge base and lifts matching accuracy from 74.5% to 90.2% across 102 real-world test queries. The per-task overhead is 0.004 ms, and regex responses are 670 to 63,300 times faster than large language model inference. The architecture is application and logic decoupled and deployable in Model Context Protocol-compatible systems.

Authors

Institutions

Publication Details

Journal
Electronics
Published
2026-09-14
DOI
https://doi.org/10.3390/electronics15184166
Primary Topic
Software System Performance and Reliability
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

ReliHarness: A Self-Learning Reliability Harness for LLM Agent Tool Execution

Wendi Feng, Chuanchang Liu, W. Chen, Libo Cao et al.
Electronics
Software System Performance and Reliability
article

ReliHarness: A Self-Learning Reliability Harness for LLM Agent Tool Execution

Wendi Feng, Chuanchang Liu, W. Chen, Libo Cao, Bo Cheng, Penghan Song, Qichao Lu
article en

Abstract

Large language model agents invoke external tools through the Model Context Protocol to perform actions beyond text generation. Once a background thread is dispatched, even the latest Model Context Protocol Tasks extension exposes only coarse-grained, self-reported task states, so a thread that deadlocks or crashes silently remains reported as working and invisible to the client, which is not acceptable in some industrial applications that need long monitoring. We present ReliHarness, a self-learning dual-channel execution-layer reliability harness framework that instead probes each worker thread at a lower level through two independent checks: heartbeat timeout detection and operating system-level thread liveness detection. The proposed framework contributes a five-state lifecycle machine, dual-channel detection, bi-layer evidence logging, a four-stage cascade pipeline, and a self-learning command-matching module. Experiments on a Phytium S5000C and Ascend 310P platform achieve 100% detection coverage for the tested crash and deadlock scenarios with zero false positives. The self-learning module expands the knowledge base and lifts matching accuracy from 74.5% to 90.2% across 102 real-world test queries. The per-task overhead is 0.004 ms, and regex responses are 670 to 63,300 times faster than large language model inference. The architecture is application and logic decoupled and deployable in Model Context Protocol-compatible systems.

ElectronicsVol. 15(18)
Beijing University of Posts and Telecommunications (CN), China Academy of Information and Communications Technology (CN), Beijing Information Science & Technology University (CN)
Openalex Percentile: Top 8%
Software System Performance and Reliability
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.