Beyond accuracy: resource-aware evaluation of deep learning models for remote sensing image classification
Accurate and computationally efficient scene recognition is necessary for deploying remote sensing models in resource-constrained settings. Convolutional and transformer-based encoders have achieved strong performance on aerial scene benchmarks, yet systematic comparisons regarding the computational cost across architectures are limited, leaving practitioners without clear guidance for deployment-aware model selection. In this paper, we benchmark nine representative deep learning encoders spanning the major architectural families under a deployment-aware protocol on four remote sensing datasets: UC Merced Land Use, Aerial Image Dataset (AID), remote sensing image scene classification (RESISC)45, and Multi Label Remote Sensing dataset (MLRSNet). Models are evaluated on standard classification metrics alongside parameters, floating-point operations, and inference latency. To support accuracy–efficiency comparison, we introduce the Predictive Model Efficiency (PME), a deployment-aware, resource-normalized metric that combines Macro-F1 with log-normalized computational penalties within a bounded, interpretable range. PME’s novelty lies in its purpose-driven integration of Macro-F1 with both theoretical (FLOPs, parameters) and empirical (latency) efficiency measures, a bounded sigmoid transformation for cross-dataset comparability, and configurable, task-dependent weighting that enables adaptation to different deployment scenarios. All results are scoped to transfer learning and apply specifically to deployment settings where full model training is computationally prohibitive. Under our transfer-learning protocol—with fixed pretrained backbones and limited fine-tuning—convolutional and hybrid architectures consistently achieve better accuracy–efficiency trade-offs, while attention-based encoders incur higher training costs for modest accuracy gains. Peak F1 scores reached 98% (UC Merced), 94% (AID), 99% (RESISC45), and 98% (MLRSNet). While transformers matched or outperformed convolutional neural network/hybrids on smaller datasets, EdgeNeXt-s achieved the highest accuracy on MLRSNet (98%), outperforming all transformers by 2% points while incurring substantially a lower computational cost. These findings confirm prior studies, show the trade-offs between performance and efficiency across models, and highlight the importance of selecting architectures that balance performance with available hardware resources. Our codes are publicly available at https://github.com/qaixerabbas/resisc-benchmark
Authors
- Hamza Sajjad (ORCID: https://orcid.org/0000-0002-6954-938X)
- Muhammad Irzam Liaqat (ORCID: https://orcid.org/0009-0009-1265-7337)
- Qaiser Abbas (ORCID: https://orcid.org/0000-0003-1870-0884)
- Farid Ud Din Masood Khan
- Muhammad Laeeq uz Zaman Khan
Institutions
- IMT School for Advanced Studies Lucca (IT)
- University of Engineering and Technology Lahore (PK)
- Hamad bin Khalifa University (QA)
- Sasol (Germany) (DE)
Publication Details
- Journal
- International Journal of Remote Sensing
- Published
- 2026-09-24
- DOI
- https://doi.org/10.1080/01431161.2026.2736868
- Primary Topic
- Remote-Sensing Image Classification
- Type
- article
- Field-Weighted Citation Impact
- 0.00