Machine Learning for Autism Spectrum Disorder Prediction: A Review of Data Augmentation and Feature Selection Techniques

Autism spectrum disorder (ASD) is a complex neurodevelopmental condition characterized by persistent difficulties in social communication, social interaction, and repetitive behaviors. Early and accurate diagnosis is essential but is often hindered by subjective clinical assessments, limited data availability, and inconsistencies in existing diagnostic tools. This review evaluates the role of machine learning and deep learning approaches in improving ASD prediction, with a particular focus on two important yet relatively underexplored methodological components: data augmentation and feature selection. A structured literature search was conducted across major scientific databases, including IEEE Xplore, PubMed, Scopus, and Google Scholar, to identify studies published between 2021 and 2024. The methodological quality and risk of bias of the included studies were assessed using the Prediction Model Risk of Bias Assessment Tool. A total of 26 peer-reviewed studies were included based on their relevance to machine learning/deep learning-based ASD prediction and their explicit application of data augmentation or feature selection techniques. Data augmentation methods were categorized into conventional approaches, such as geometric and color-space transformations, and advanced techniques, including generative adversarial network-based synthetic data generation. Although augmentation techniques may improve model robustness and help address dataset scarcity, relatively few studies conducted ablation analyses to isolate the contribution of individual augmentation strategies. Feature selection approaches were classified into filter, wrapper, and embedded methods. Commonly used techniques included information gain, chi-square tests, recursive feature elimination, and elastic net regularization. While these methods may improve predictive performance and model interpretability, they are frequently applied without sufficient empirical justification or biological interpretation. Overall, this review highlights important methodological limitations, including limited external validation, insufficient ablation analyses, and inadequate evaluation frameworks, which reduce confidence in reported performance improvements and model generalizability. Future research should emphasize methodological transparency, robust validation strategies, multimodal data integration, and clinically interpretable modeling approaches to improve the reliability and clinical applicability of ASD prediction systems.

Authors

Institutions

Publication Details

Journal
Health care science
Published
2026-09-07
DOI
https://doi.org/10.1002/hcs2.70098
Primary Topic
Autism Spectrum Disorder Research
Type
article
Field-Weighted Citation Impact
0.00
Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Machine Learning for Autism Spectrum Disorder Prediction: A Review of Data Augmentation and Feature Selection Techniques

Feng Dong, Sahar Ahmed M Alkhaibari
Health care science
Autism Spectrum Disorder Research
article

Machine Learning for Autism Spectrum Disorder Prediction: A Review of Data Augmentation and Feature Selection Techniques

Feng Dong, Sahar Ahmed M Alkhaibari
article en

Abstract

Autism spectrum disorder (ASD) is a complex neurodevelopmental condition characterized by persistent difficulties in social communication, social interaction, and repetitive behaviors. Early and accurate diagnosis is essential but is often hindered by subjective clinical assessments, limited data availability, and inconsistencies in existing diagnostic tools. This review evaluates the role of machine learning and deep learning approaches in improving ASD prediction, with a particular focus on two important yet relatively underexplored methodological components: data augmentation and feature selection. A structured literature search was conducted across major scientific databases, including IEEE Xplore, PubMed, Scopus, and Google Scholar, to identify studies published between 2021 and 2024. The methodological quality and risk of bias of the included studies were assessed using the Prediction Model Risk of Bias Assessment Tool. A total of 26 peer-reviewed studies were included based on their relevance to machine learning/deep learning-based ASD prediction and their explicit application of data augmentation or feature selection techniques. Data augmentation methods were categorized into conventional approaches, such as geometric and color-space transformations, and advanced techniques, including generative adversarial network-based synthetic data generation. Although augmentation techniques may improve model robustness and help address dataset scarcity, relatively few studies conducted ablation analyses to isolate the contribution of individual augmentation strategies. Feature selection approaches were classified into filter, wrapper, and embedded methods. Commonly used techniques included information gain, chi-square tests, recursive feature elimination, and elastic net regularization. While these methods may improve predictive performance and model interpretability, they are frequently applied without sufficient empirical justification or biological interpretation. Overall, this review highlights important methodological limitations, including limited external validation, insufficient ablation analyses, and inadequate evaluation frameworks, which reduce confidence in reported performance improvements and model generalizability. Future research should emphasize methodological transparency, robust validation strategies, multimodal data integration, and clinically interpretable modeling approaches to improve the reliability and clinical applicability of ASD prediction systems.

Health care science
University of Strathclyde (GB), University of Tabuk (SA)
Openalex Percentile: Top 27%
Autism Spectrum Disorder Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.