Fuzzy k-Nearest Neighbor Classifier Based on Ordered Fuzzy Numbers for Early Fault Detection in Grid-Connected Photovoltaic Systems
The growing scale of photovoltaic (PV) installations creates an urgent need for automated fault detection systems that operate in real time on resource-constrained monitoring hardware. Although deep learning methods achieve high classification accuracy on PV fault benchmarks, their computational requirements make them impractical for deployment on embedded devices such as smart inverters and industrial IoT controllers. This paper proposes a lightweight classification approach based on the fuzzy k-Nearest Neighbor (Fuzzy kNN) algorithm, in which every electrical measurement is represented as an Ordered Fuzzy Number (OFN) whose spread is automatically calibrated from the local standard deviation of the measurement window. This adaptive fuzzification encodes the inherent sensor noise and environmental variability of SCADA measurements without any manual parameter tuning. Six fuzzy distance metrics, obtained by combining three defuzzification operators (FOM, LOM, MOM) with the Euclidean and Manhattan distance functions, were evaluated on the public GPVS-Faults benchmark containing approximately 1.8 million samples describing seven fault types in a grid-connected PV system operating under MPPT and IPPT control. In binary anomaly detection, the proposed method achieved an accuracy of 92.97% ± 1.17% (k = 3, MOM defuzzification with Manhattan distance), which is statistically comparable to the Random Forest baseline (92.42% ± 0.65%) while offering approximately one hundred times faster inference (1.2 ms versus 120 ms per sample) and a model footprint of only 2 MB. In multiclass fault type classification, the method reached 96.18% ± 1.80% accuracy against 97.88% ± 0.69% for Random Forest. A consistent and previously unreported observation is that the Manhattan distance systematically outperforms the Euclidean distance on three-phase electrical measurements, improving accuracy by approximately 0.90 percentage points across all tested configurations. The complete source code and the experimental pipeline are publicly released to ensure full reproducibility of the reported results. To assess generalization rigorously, a stratified group cross-validation was additionally performed in which all windows from a given experimental recording are confined to a single fold; under this leakage-free protocol every evaluated method, including Random Forest, degrades to the 50–62% range, which shows that cross-recording transfer is an intrinsic difficulty of the single-run GPVS-Faults benchmark rather than a weakness specific to the proposed classifier.
Authors
- Łukasz Apiecionek (ORCID: https://orcid.org/0000-0003-4565-5505)
Institutions
- Kazimierz Wielki University in Bydgoszcz (PL)
Publication Details
- Journal
- Energies
- Published
- 2026-09-17
- DOI
- https://doi.org/10.3390/en19184395
- Primary Topic
- Photovoltaic System Optimization Techniques
- Type
- article
- Field-Weighted Citation Impact
- 0.00
Funders
- Ministry of Science and Higher Education of the Russian Federation