Automated screening of autism via a pretrained vision-language model in naturalistic caregiver-child interactions

Early autism screening is essential for timely intervention but remains constrained by specialist dependence, costly protocols, and limited accessibility. We present Autism-CLIP, a vision-language model that automates screening from 3-min naturalistic caregiver-child interaction videos recorded in toddlers aged 15–24 months. By integrating contrastive learning, large language model–generated supervision, domain-specific architecture, and hybrid training, Autism-CLIP aligns visual behavioral cues with autism-related assessment outcomes, enabling annotation-free inference. In multicenter evaluation, Autism-CLIP achieved an AUC of 0.903 in the internal cohort (n = 103) and 0.873 in the external cohort (n = 102), outperforming conventional machine learning methods. Attention mapping highlighted clinically relevant markers, including gaze aversion and limited social reciprocity. Requiring only smartphone-captured videos without manual coding or controlled settings, Autism-CLIP offers a scalable, low-cost approach to broaden access to early autism screening and support prioritization of clinical resources. Here, the authors develop Autism-CLIP, a vision–language model for automated early autism screening using 3-min home videos of toddlers aged 15–24 months.

Authors

Institutions

Publication Details

Journal
Nature Communications
Published
2026-09-16
DOI
https://doi.org/10.1038/s41467-026-77721-8
Primary Topic
Autism Spectrum Disorder Research
Type
article
Field-Weighted Citation Impact
0.00

Funders

Controls
|||
ALL TIME
JAN
FEB
MAR
APR
MAY
JUN
JUL
AUG
SEP
article

Automated screening of autism via a pretrained vision-language model in naturalistic caregiver-child interactions

Huishi Huang, Fei Wu, Hongzhu Deng, Shaoli Lv et al.
Nature Communications
Autism Spectrum Disorder Research
article

Automated screening of autism via a pretrained vision-language model in naturalistic caregiver-child interactions

Huishi Huang, Fei Wu, Hongzhu Deng, Shaoli Lv, Yu Xing, Zhengxing Huang, Yijie Li, Cong You, Ruikang Deng
article en

Abstract

Early autism screening is essential for timely intervention but remains constrained by specialist dependence, costly protocols, and limited accessibility. We present Autism-CLIP, a vision-language model that automates screening from 3-min naturalistic caregiver-child interaction videos recorded in toddlers aged 15–24 months. By integrating contrastive learning, large language model–generated supervision, domain-specific architecture, and hybrid training, Autism-CLIP aligns visual behavioral cues with autism-related assessment outcomes, enabling annotation-free inference. In multicenter evaluation, Autism-CLIP achieved an AUC of 0.903 in the internal cohort (n = 103) and 0.873 in the external cohort (n = 102), outperforming conventional machine learning methods. Attention mapping highlighted clinically relevant markers, including gaze aversion and limited social reciprocity. Requiring only smartphone-captured videos without manual coding or controlled settings, Autism-CLIP offers a scalable, low-cost approach to broaden access to early autism screening and support prioritization of clinical resources. Here, the authors develop Autism-CLIP, a vision–language model for automated early autism screening using 3-min home videos of toddlers aged 15–24 months.

Nature Communications
Sun Yat-sen University (CN), Zhejiang University of Science and Technology (CN), Shanghai Jiao Tong University (CN), Third Affiliated Hospital of Sun Yat-sen University (CN)
National Natural Science Foundation of China, Department of Science and Technology for Social Development, Guangzhou Municipal Science and Technology Project, National Key Research and Development Program of China, Natural Science Foundation of Zhejiang Province
Gender equality
Openalex Percentile: Top 10%
Autism Spectrum Disorder Research
AI Navigator

Ask Laika to Summarize, Analyze, and Connect papers live on the map.

Summarize Papers & Methodologies

Extract key findings, datasets, and comparative methods across publications.

Benchmark Rankings & Visual Analytics

Rank top research institutions, authors, funders, topics, and journals by Field-Weighted Citation Impact (FWCI) and paper volume with instant charts.

Connect Distant Disciplines

Bridge topological clusters on the map to find hidden collaborative intersections.