Towards the development of explainable machine learning models to recognize the faces of autistic children: a brief report
Bibliographic record
Abstract
Purpose Machine learning with image classification has shown promise in supporting the detection of autism in children, but the development of explainable models is still lacking. To address this issue, the purpose of this study was to compare the development of explainable models using two different algorithms to identify the facial features that deep neural networks used to classify children as autistic or non-autistic. Design/methodology/approach First, this paper trained and tested different models on the Autistic Children Facial Image Data Set and selected the one that produced the highest accuracy. Following the identification of the best model, the analyses compared two methods to examine explainability: Local Interpretable Model-agnostic Explanations and Randomized Input Sampling for Explanation of black-box models. Findings Overall, the best model, ViT_Huge_14, produced an accuracy of 92%. Moreover, Local Interpretable Model-agnostic Explanations resulted in more explainable models than Randomized Input Sampling for Explanation of black-box models. Albeit promising, researchers must conduct further studies to examine the generalizability of the results and consider ethical issues before recommending facial image classification as a component of a multimethod approach to screening and diagnosis. Originality/value To the best of the authors’ knowledge, this study is the first to examine the development of explainable models to detect autism using facial features.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.014 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".