Autism Diagnostic Assessments for Gender-Diverse Individuals: A Modified Delphi Study of Clinician Experts in the Fields of Autism and Gender Diversity
Bibliographic record
Abstract
Objective There is limited evidence-based guidance about how autism diagnostic assessments should be conducted for gender-diverse people. We aimed to integrate expert knowledge on key clinical considerations for these assessments. Method We conducted a modified Delphi study. World experts in the field ( N =21) were invited to complete two rounds of surveys. Survey One collected open-text responses about key clinical considerations when conducting diagnostic assessments, structured around the DSM-5-TR criteria for autism, across age-ranges. Experts were asked to rate the importance of each consideration they listed. A content analysis was conducted to synthesise and collate similar considerations, alongside descriptive statistics of importance ratings. Survey Two presented the resulting considerations and mean importance ratings, with experts re-rating their importance. Statements rated as at least ‘important' and that had a standard deviation of less than 1.0 were reported. Results Round one resulted in 65 individual statements, of which 37 met our definition for reporting. These statements, summarizing expert opinions, were categorised as being (1) general considerations for assessments, (2) linked to the DSM-5-TR autism criteria (A-E), or (3) practical considerations for working with the gender-diverse population. They highlighted areas to be considered during assessments, such as ways in which the features of autism may intersect with gender diversity, and practical considerations for increasing comfort and engagement of gender-diverse individuals undergoing an autism assessment. Conclusion The summary of expert opinions provides preliminary considerations for clinicians working in this field, and for researchers to use as hypotheses for empirical investigations.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".