Beyond accuracy: resource-aware evaluation of deep learning models for remote sensing image classification

Accurate and computationally efficient scene recognition is necessary for deploying remote sensing models in resource-constrained settings. Convolutional and transformer-based encoders have achieved strong performance on aerial scene benchmarks, yet systematic comparisons regarding the computational cost across architectures are limited, leaving practitioners without clear guidance for deployment-aware model selection. In this paper, we benchmark nine representative deep learning encoders spanning the major architectural families under a deployment-aware protocol on four remote sensing datasets: UC Merced Land Use, Aerial Image Dataset (AID), remote sensing image scene classification (RESISC)45, and Multi Label Remote Sensing dataset (MLRSNet). Models are evaluated on standard classification metrics alongside parameters, floating-point operations, and inference latency. To support accuracy–efficiency comparison, we introduce the Predictive Model Efficiency (PME), a deployment-aware, resource-normalized metric that combines Macro-F1 with log-normalized computational penalties within a bounded, interpretable range. PME’s novelty lies in its purpose-driven integration of Macro-F1 with both theoretical (FLOPs, parameters) and empirical (latency) efficiency measures, a bounded sigmoid transformation for cross-dataset comparability, and configurable, task-dependent weighting that enables adaptation to different deployment scenarios. All results are scoped to transfer learning and apply specifically to deployment settings where full model training is computationally prohibitive. Under our transfer-learning protocol—with fixed pretrained backbones and limited fine-tuning—convolutional and hybrid architectures consistently achieve better accuracy–efficiency trade-offs, while attention-based encoders incur higher training costs for modest accuracy gains. Peak F1 scores reached 98% (UC Merced), 94% (AID), 99% (RESISC45), and 98% (MLRSNet). While transformers matched or outperformed convolutional neural network/hybrids on smaller datasets, EdgeNeXt-s achieved the highest accuracy on MLRSNet (98%), outperforming all transformers by 2% points while incurring substantially a lower computational cost. These findings confirm prior studies, show the trade-offs between performance and efficiency across models, and highlight the importance of selecting architectures that balance performance with available hardware resources. Our codes are publicly available at https://github.com/qaixerabbas/resisc-benchmark

Authors

Institutions

Publication Details

Journal
International Journal of Remote Sensing
Published
2026-09-24
DOI
https://doi.org/10.1080/01431161.2026.2736868
Primary Topic
Remote-Sensing Image Classification
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Beyond accuracy: resource-aware evaluation of deep learning models for remote sensing image classification

Hamza Sajjad, Muhammad Irzam Liaqat, Qaiser Abbas, Farid Ud Din Masood Khan et al.
International Journal of Remote Sensing
Remote-Sensing Image Classification
article

Beyond accuracy: resource-aware evaluation of deep learning models for remote sensing image classification

Hamza Sajjad, Muhammad Irzam Liaqat, Qaiser Abbas, Farid Ud Din Masood Khan, Muhammad Laeeq uz Zaman Khan
article en

Abstract

Accurate and computationally efficient scene recognition is necessary for deploying remote sensing models in resource-constrained settings. Convolutional and transformer-based encoders have achieved strong performance on aerial scene benchmarks, yet systematic comparisons regarding the computational cost across architectures are limited, leaving practitioners without clear guidance for deployment-aware model selection. In this paper, we benchmark nine representative deep learning encoders spanning the major architectural families under a deployment-aware protocol on four remote sensing datasets: UC Merced Land Use, Aerial Image Dataset (AID), remote sensing image scene classification (RESISC)45, and Multi Label Remote Sensing dataset (MLRSNet). Models are evaluated on standard classification metrics alongside parameters, floating-point operations, and inference latency. To support accuracy–efficiency comparison, we introduce the Predictive Model Efficiency (PME), a deployment-aware, resource-normalized metric that combines Macro-F1 with log-normalized computational penalties within a bounded, interpretable range. PME’s novelty lies in its purpose-driven integration of Macro-F1 with both theoretical (FLOPs, parameters) and empirical (latency) efficiency measures, a bounded sigmoid transformation for cross-dataset comparability, and configurable, task-dependent weighting that enables adaptation to different deployment scenarios. All results are scoped to transfer learning and apply specifically to deployment settings where full model training is computationally prohibitive. Under our transfer-learning protocol—with fixed pretrained backbones and limited fine-tuning—convolutional and hybrid architectures consistently achieve better accuracy–efficiency trade-offs, while attention-based encoders incur higher training costs for modest accuracy gains. Peak F1 scores reached 98% (UC Merced), 94% (AID), 99% (RESISC45), and 98% (MLRSNet). While transformers matched or outperformed convolutional neural network/hybrids on smaller datasets, EdgeNeXt-s achieved the highest accuracy on MLRSNet (98%), outperforming all transformers by 2% points while incurring substantially a lower computational cost. These findings confirm prior studies, show the trade-offs between performance and efficiency across models, and highlight the importance of selecting architectures that balance performance with available hardware resources. Our codes are publicly available at https://github.com/qaixerabbas/resisc-benchmark

International Journal of Remote Sensing
IMT School for Advanced Studies Lucca (IT), University of Engineering and Technology Lahore (PK), Hamad bin Khalifa University (QA), Sasol (Germany) (DE)
Openalex Percentile: Top 14%
Remote-Sensing Image Classification
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.