When Data Augmentation Falls Short: Wi-Fi Fingerprint-Based Indoor Localization Revisited

Generative data augmentation has been widely explored in Wi-Fi fingerprint-based indoor localization to reduce the cost of dense radiomap construction, with many studies reporting substantial localization improvements. However, these comparisons typically rely on default or weakly optimized baseline regressors, making it difficult to determine whether reported gains reflect genuine synthesis quality or merely compensate for suboptimal baselines. In this paper, we systematically investigate under which conditions generative augmentation is actually justified. We fine-tune four widely used localization regressors—kNN, SVR, XGBoost, and DNN—using Bayesian hyperparameter optimization and establish strong non-augmented baselines across radiomaps with controlled levels of spatial sparsity, constructed via farthest-point sampling. We then train five representative generative models—VAE, GAN, DDPM, DiT, and TDPM—within a unified augmentation pipeline that includes quality filtering and pseudo-labeling, and benchmark them against these baselines. Using two publicly available datasets, we show that none of the generative models consistently outperforms a non-augmented, fine-tuned baseline regressor such as XGBoost or kNN, across a wide range of sparsity levels. We further show that these conclusions are robust to three potential confounders: various proportions of synthetic data, the choice of localization regressor (ruling out circularity with the pseudo-labeling model), and the dataset itself, since the findings on the first dataset replicate on a second, structurally different building. These findings suggest that reported augmentation benefits in prior work may partly reflect under-optimized baselines rather than genuine synthesis quality, and that generative augmentation should be treated as a conditional last resort rather than a universal improvement strategy.

Authors

Institutions

Publication Details

Journal
Sensors
Published
2026-08-26
DOI
https://doi.org/10.3390/s26175392
Primary Topic
Indoor and Outdoor Localization Technologies
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

When Data Augmentation Falls Short: Wi-Fi Fingerprint-Based Indoor Localization Revisited

Marko Ristin, Shinnazar Seytnazarov, Nurbek Malikov
Sensors
Indoor and Outdoor Localization Technologies
article

When Data Augmentation Falls Short: Wi-Fi Fingerprint-Based Indoor Localization Revisited

Marko Ristin, Shinnazar Seytnazarov, Nurbek Malikov
article en

Abstract

Generative data augmentation has been widely explored in Wi-Fi fingerprint-based indoor localization to reduce the cost of dense radiomap construction, with many studies reporting substantial localization improvements. However, these comparisons typically rely on default or weakly optimized baseline regressors, making it difficult to determine whether reported gains reflect genuine synthesis quality or merely compensate for suboptimal baselines. In this paper, we systematically investigate under which conditions generative augmentation is actually justified. We fine-tune four widely used localization regressors—kNN, SVR, XGBoost, and DNN—using Bayesian hyperparameter optimization and establish strong non-augmented baselines across radiomaps with controlled levels of spatial sparsity, constructed via farthest-point sampling. We then train five representative generative models—VAE, GAN, DDPM, DiT, and TDPM—within a unified augmentation pipeline that includes quality filtering and pseudo-labeling, and benchmark them against these baselines. Using two publicly available datasets, we show that none of the generative models consistently outperforms a non-augmented, fine-tuned baseline regressor such as XGBoost or kNN, across a wide range of sparsity levels. We further show that these conclusions are robust to three potential confounders: various proportions of synthetic data, the choice of localization regressor (ruling out circularity with the pseudo-labeling model), and the dataset itself, since the findings on the first dataset replicate on a second, structurally different building. These findings suggest that reported augmentation benefits in prior work may partly reflect under-optimized baselines rather than genuine synthesis quality, and that generative augmentation should be treated as a conditional last resort rather than a universal improvement strategy.

SensorsVol. 26(17)
ZHAW Zurich University of Applied Sciences (CH), Nazarbayev University (KZ)
Nazarbayev University
Industry, innovation and infrastructure
Openalex Percentile: Top 19%
Indoor and Outdoor Localization Technologies
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.