eXclusionarY: Ten years later, where are the sex chromosomes in GWAS?
Bibliographic record
Abstract
Summary Ten years ago, a detailed analysis of genome-wide association studies showed that only 33% of the studies included the X chromosome. Multiple recommendations were made to combat eXclusion. Here we re-surveyed the research landscape to determine if these earlier recommendations had been translated. Unfortunately, among the summary statistics reported in 2021 in the NHGRI-EBI GWAS catalog, only 25% provided results for the X chromosome and 3% for the Y chromosome, suggesting that the eXclusion phenomenon documented earlier not only persists but has also expanded into an eXclusionarY problem. Normalizing by physical length of the chromosome, the average number of studies published until 11/29/22 with genome-wide significant findings on the X chromosome is ~1 study/Mb. In contrast, it ranges from ~6 to ~16 studies/Mb for chromosomes 4 and 19, respectively. Compared with the autosomal growth rate of ~0.086 studies/Mb/year over the last decade, studies of the X chromosome grew at less than one-seventh that rate, only ~0.012 studies/Mb/year. Among the studies that reported significant association on the X chromosome, there were extreme heterogeneities in how they analyzed the data and documented the results, suggesting the need for guidelines. Not surprisingly, among the 430 scores sampled from the PolyGenic Score catalog, 0% contained weights for sex chromosomal SNPs. To overcome the dearth of sex chromosome analyses, we provide five sets of recommendations and future directions. Finally, until the sex chromosomes are included in a whole-genome study, instead of GWAS, we propose they be more properly referred to as “AWAS” for “autosome-wide scan”.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.085 | 0.219 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.002 | 0.006 |
| Scholarly communication | 0.008 | 0.009 |
| Open science | 0.003 | 0.003 |
| Research integrity | 0.004 | 0.015 |
| Insufficient payload (model declined to judge) | 0.012 | 0.004 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".