Biallelic variants identified in 36 Pakistani families and trios with autism spectrum disorder
Bibliographic record
Abstract
With its high rate of consanguineous marriages and diverse ethnic population, little is currently understood about the genetic architecture of autism spectrum disorder (ASD) in Pakistan. Pakistan has a highly ethnically diverse population, yet with a high proportion of endogamous marriages, and is therefore anticipated to be enriched for biallelic disease-relate variants. Here, we attempt to determine the underlying genetic abnormalities causing ASD in thirty-six small simplex or multiplex families from Pakistan. Microarray genotyping followed by homozygosity mapping, copy number variation analysis, and whole exome sequencing were used to identify candidate. Given the high levels of consanguineous marriages among these families, autosomal recessively inherited variants were prioritized, however de novo/dominant and X-linked variants were also identified. The selected variants were validated using Sanger sequencing. Here we report the identification of sixteen rare or novel coding variants in fifteen genes (ARAP1, CDKL5, CSMD2, EFCAB12, EIF3H, GML, NEDD4, PDZD4, POLR3G, SLC35A2, TMEM214, TMEM232, TRANK1, TTC19, and ZNF292) in affected members in eight of the families, including ten homozygous variants in four families (nine missense, one loss of function). Three heterozygous de novo mutations were also identified (in ARAP1, CSMD2, and NEDD4), and variants in known X-linked neurodevelopmental disorder genes CDKL5 and SLC35A2. The current study offers information on the genetic variability associated with ASD in Pakistan, and demonstrates a marked enrichment for biallelic variants over that reported in outbreeding populations. This information will be useful for improving approaches for studying ASD in populations where endogamy is commonly practiced.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".