AI-powered facial analysis accelerates autism screening through virtual reality attention tasks

Researchers have developed a dual-branch deep-learning system that uses facial geometry and texture analysis during virtual-reality attention tasks to provide a faster, less subjective autism screening tool applicable across diverse populations.

Researchers have developed a dual-branch deep-learning system that looks at children’s facial geometry and texture while they carry out attention tasks in a virtual-reality classroom, in an attempt to make autism screening faster and less subjective. According to a study in Machine Learning with Applications, the strongest version of the model, called GDFN, performed best on a webcam-based dataset of school-age children from 15 nationalities, while explainable AI tools helped show which facial regions influenced its decisions.

The work targets a longstanding problem in autism care: assessment is still heavily dependent on specialist observation, structured interviews and rating scales, which can take time and vary between clinicians. The World Health Organization estimates that about one in 100 children worldwide are diagnosed with autism, and the US Centres for Disease Control and Prevention puts the figure at one in 36 in the United States. That gap between need and diagnosis matters, especially because delayed identification can mean children miss out on early support during crucial developmental years.

The idea of reading autism-related clues from faces is not new. Earlier studies, including work published in PMC and on ScienceDirect, found that deep-learning systems can identify facial patterns associated with autism, and that combining image data with facial landmarks can improve performance and interpretability. The new study builds on that approach by pairing learned image features with clinically derived measurements, including 31 distances between facial landmarks, rather than relying on image analysis alone.

The researchers also tested the system in a realistic setting: children were recorded by webcam during a virtual-reality continuous performance test, an attention task in which they respond to target stimuli and ignore distractions. The dataset included 85 children aged seven to 12, with both autistic and typically developing participants, and it was reviewed under ethics approval with parental consent and child assent. The study also used a separate public facial-image dataset to check whether the model could generalise beyond the original group.

The results suggest that hybrid systems work better than single-feature models. An ablation analysis showed that geometric measurements, SIFT-based texture descriptors and deep features each contributed differently, and that combining them improved screening accuracy. The authors also used Grad-CAM and saliency mapping to show where the model was looking, an important step for clinical trust. Even so, the study presents the tool as a screening aid, not a diagnosis, and the authors say the next challenge is broader testing across populations and ethnic groups.

Disclaimer: This content is for informational purposes only and is not intended to be a substitute for professional medical judgment, advice, diagnosis, or treatment.