Characterising On-Device AI Failure Modes Under Joint Fault Injection and Runtime Observability

Deploying artificial intelligence models directly on edge devices exposes machine learning runtimes to hardware volatile environments, including thermal limits, volatile system memory pressure, and corrupted input data. While previous work evaluates software-level fault injection and runtime observability in isolation, no empirical study has applied fault injection and high-frequency device tracing in a combined pipeline across heterogeneous model architectures. We present a joint empirical study combining SATE AI (software-level fault injection) and EdgePulse (non-intrusive runtime observability) on a physical Android handset (Tecno CH7n, MediaTek Helio G35, Android 12). Across a benchmark suite of 280 real-device inference traces covering three model runtimes (TensorFlow Lite, ONNX Runtime, and llama.cpp GGUF) and four execution scenarios (baseline, memory pressure, malformed input, and thermal stress), we demonstrate three principal findings: 1. Thermal stress causes extreme latency degradation (+486.8% for MobileNetV3 TFLite, +406.6% for ResNet-18 ONNX), while Android's high-level PowerManager.currentThermalStatus API reported nominal across 100% of the 70 thermal stress runs, replicating and reinforcing the thermal API blind spot.2. Malformed input injection produces a distinct bimodal failure signature in quantized LLMs: 80% of runs complete within nominal latency (807-2680 ms), while 20% enter a 20-second pathological generation loop (19,702 ms and 19,963 ms), creating an extreme standard deviation (std = 7,244.2 ms).3. Vision models and Large Language Models exhibit inverted fault tolerance profiles: vision models are thermal-sensitive but malformed-input-resilient, whereas edge LLMs are malformed-input-sensitive but less thermal-sensitive. Both frameworks are open-source software under the MIT license, available at sate_ai and edgepulse.

Authors

Publication Details

Journal
Zenodo (CERN European Organization for Nuclear Research)
Published
2026-10-09
DOI
https://doi.org/10.5281/zenodo.23262269
Primary Topic
Software System Performance and Reliability
Type
preprint
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
OCT
preprint

Characterising On-Device AI Failure Modes Under Joint Fault Injection and Runtime Observability

Muhammad Assad Ullah
Zenodo (CERN European Organization for Nuclear Research)
Software System Performance and Reliability
preprint

Characterising On-Device AI Failure Modes Under Joint Fault Injection and Runtime Observability

Muhammad Assad Ullah
preprint en

Abstract

Deploying artificial intelligence models directly on edge devices exposes machine learning runtimes to hardware volatile environments, including thermal limits, volatile system memory pressure, and corrupted input data. While previous work evaluates software-level fault injection and runtime observability in isolation, no empirical study has applied fault injection and high-frequency device tracing in a combined pipeline across heterogeneous model architectures. We present a joint empirical study combining SATE AI (software-level fault injection) and EdgePulse (non-intrusive runtime observability) on a physical Android handset (Tecno CH7n, MediaTek Helio G35, Android 12). Across a benchmark suite of 280 real-device inference traces covering three model runtimes (TensorFlow Lite, ONNX Runtime, and llama.cpp GGUF) and four execution scenarios (baseline, memory pressure, malformed input, and thermal stress), we demonstrate three principal findings: 1. Thermal stress causes extreme latency degradation (+486.8% for MobileNetV3 TFLite, +406.6% for ResNet-18 ONNX), while Android's high-level PowerManager.currentThermalStatus API reported nominal across 100% of the 70 thermal stress runs, replicating and reinforcing the thermal API blind spot.2. Malformed input injection produces a distinct bimodal failure signature in quantized LLMs: 80% of runs complete within nominal latency (807-2680 ms), while 20% enter a 20-second pathological generation loop (19,702 ms and 19,963 ms), creating an extreme standard deviation (std = 7,244.2 ms).3. Vision models and Large Language Models exhibit inverted fault tolerance profiles: vision models are thermal-sensitive but malformed-input-resilient, whereas edge LLMs are malformed-input-sensitive but less thermal-sensitive. Both frameworks are open-source software under the MIT license, available at sate_ai and edgepulse.

Zenodo (CERN European Organization for Nuclear Research)
Software System Performance and Reliability
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.