Deep Learning Models for IoT: MobileNetV2 Compression Using Structured Pruning and Post-Training Quantization
Deep learning models are needed to develop light and efficient algorithms because of fast developments on embedded devices in the case of internet of things (IoT). The traditional convolutional neural networks are accurate but are not successful enough to support edge deployment due to missing resources in these devices. A better version of MobileNetV2 has also been realized by incorporating two advanced compression procedures such as structured pruning and post-training quantization. This project is based on the principles spelled out by Zhang and Li (2023) of a Review of artificial intelligence in embedded systems to transition into the practical implementation systems. The CIFAR-10 dataset was used in order to deliver a benchmark of the models performance in terms of accuracy and deployment metrics, such as size and latency speed. Research measurements have proven that with the use of pruning techniques and quantization techniques modeling size can reduce to a maximum of 80 percent without reducing the results in the prediction process. Embedded devices, including microcontrollers and edge AI boards, are some of the hardware entities explored in this paper in terms of the specifications and deployment abilities. The document rounds off with research recommendations on the use of embedded AIs and an illustration and review comparison of the trade-offs.
Authors
- Fakhira Afzal
- Usman Ashraf (ORCID: https://orcid.org/0000-0001-6288-3513)
- Muhammad Ahsan Ali
- Hafiz Muhammad Yousha
- Abdullah Qamar
- Shanza Riaz
- Hamza Afzal
Institutions
- Islamia University of Bahawalpur (PK)
- Health Solutions (Sweden) (SE)
Publication Details
- Journal
- Zenodo (CERN European Organization for Nuclear Research)
- Published
- 2026-09-15
- DOI
- https://doi.org/10.5281/zenodo.22766806
- Primary Topic
- Advanced Neural Network Applications
- Type
- preprint