Fuzzy k-Nearest Neighbor Classifier Based on Ordered Fuzzy Numbers for Early Fault Detection in Grid-Connected Photovoltaic Systems

The growing scale of photovoltaic (PV) installations creates an urgent need for automated fault detection systems that operate in real time on resource-constrained monitoring hardware. Although deep learning methods achieve high classification accuracy on PV fault benchmarks, their computational requirements make them impractical for deployment on embedded devices such as smart inverters and industrial IoT controllers. This paper proposes a lightweight classification approach based on the fuzzy k-Nearest Neighbor (Fuzzy kNN) algorithm, in which every electrical measurement is represented as an Ordered Fuzzy Number (OFN) whose spread is automatically calibrated from the local standard deviation of the measurement window. This adaptive fuzzification encodes the inherent sensor noise and environmental variability of SCADA measurements without any manual parameter tuning. Six fuzzy distance metrics, obtained by combining three defuzzification operators (FOM, LOM, MOM) with the Euclidean and Manhattan distance functions, were evaluated on the public GPVS-Faults benchmark containing approximately 1.8 million samples describing seven fault types in a grid-connected PV system operating under MPPT and IPPT control. In binary anomaly detection, the proposed method achieved an accuracy of 92.97% ± 1.17% (k = 3, MOM defuzzification with Manhattan distance), which is statistically comparable to the Random Forest baseline (92.42% ± 0.65%) while offering approximately one hundred times faster inference (1.2 ms versus 120 ms per sample) and a model footprint of only 2 MB. In multiclass fault type classification, the method reached 96.18% ± 1.80% accuracy against 97.88% ± 0.69% for Random Forest. A consistent and previously unreported observation is that the Manhattan distance systematically outperforms the Euclidean distance on three-phase electrical measurements, improving accuracy by approximately 0.90 percentage points across all tested configurations. The complete source code and the experimental pipeline are publicly released to ensure full reproducibility of the reported results. To assess generalization rigorously, a stratified group cross-validation was additionally performed in which all windows from a given experimental recording are confined to a single fold; under this leakage-free protocol every evaluated method, including Random Forest, degrades to the 50–62% range, which shows that cross-recording transfer is an intrinsic difficulty of the single-run GPVS-Faults benchmark rather than a weakness specific to the proposed classifier.

Authors

Institutions

Publication Details

Journal
Energies
Published
2026-09-17
DOI
https://doi.org/10.3390/en19184395
Primary Topic
Photovoltaic System Optimization Techniques
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Fuzzy k-Nearest Neighbor Classifier Based on Ordered Fuzzy Numbers for Early Fault Detection in Grid-Connected Photovoltaic Systems

Łukasz Apiecionek
Energies
Photovoltaic System Optimization Techniques
article

Fuzzy k-Nearest Neighbor Classifier Based on Ordered Fuzzy Numbers for Early Fault Detection in Grid-Connected Photovoltaic Systems

Łukasz Apiecionek
article en

Abstract

The growing scale of photovoltaic (PV) installations creates an urgent need for automated fault detection systems that operate in real time on resource-constrained monitoring hardware. Although deep learning methods achieve high classification accuracy on PV fault benchmarks, their computational requirements make them impractical for deployment on embedded devices such as smart inverters and industrial IoT controllers. This paper proposes a lightweight classification approach based on the fuzzy k-Nearest Neighbor (Fuzzy kNN) algorithm, in which every electrical measurement is represented as an Ordered Fuzzy Number (OFN) whose spread is automatically calibrated from the local standard deviation of the measurement window. This adaptive fuzzification encodes the inherent sensor noise and environmental variability of SCADA measurements without any manual parameter tuning. Six fuzzy distance metrics, obtained by combining three defuzzification operators (FOM, LOM, MOM) with the Euclidean and Manhattan distance functions, were evaluated on the public GPVS-Faults benchmark containing approximately 1.8 million samples describing seven fault types in a grid-connected PV system operating under MPPT and IPPT control. In binary anomaly detection, the proposed method achieved an accuracy of 92.97% ± 1.17% (k = 3, MOM defuzzification with Manhattan distance), which is statistically comparable to the Random Forest baseline (92.42% ± 0.65%) while offering approximately one hundred times faster inference (1.2 ms versus 120 ms per sample) and a model footprint of only 2 MB. In multiclass fault type classification, the method reached 96.18% ± 1.80% accuracy against 97.88% ± 0.69% for Random Forest. A consistent and previously unreported observation is that the Manhattan distance systematically outperforms the Euclidean distance on three-phase electrical measurements, improving accuracy by approximately 0.90 percentage points across all tested configurations. The complete source code and the experimental pipeline are publicly released to ensure full reproducibility of the reported results. To assess generalization rigorously, a stratified group cross-validation was additionally performed in which all windows from a given experimental recording are confined to a single fold; under this leakage-free protocol every evaluated method, including Random Forest, degrades to the 50–62% range, which shows that cross-recording transfer is an intrinsic difficulty of the single-run GPVS-Faults benchmark rather than a weakness specific to the proposed classifier.

EnergiesVol. 19(18)
Kazimierz Wielki University in Bydgoszcz (PL)
Ministry of Science and Higher Education of the Russian Federation
Openalex Percentile: Top 29%
Photovoltaic System Optimization Techniques
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.