Reply to Williams et al.: Fair and Safe Eligibility Criteria for Women's Sport
Bibliographic record
Abstract
We thank Williams and colleagues [1] for recent comments reiterating our concerns about targeted sex verification based on allegation and suspicion, which motivated our initial submission [2]. It was intended as a first proposal for more ethical and equitable regulation of eligibility for women's sport, and we welcome the confirmation that several of the Williams et al. authors concur that the International Olympic Committee's (IOC's) Framework does not protect fairness for female athletes. In Lundberg et al. [3], we (including authors from Williams et al.) explained that developmental androgenization, driven by testes-derived testosterone, underpins male athletic advantage, necessitating sex-based categories in sport. We further argued that the IOC's “no presumption of advantage” [4] is logically flawed and that exclusion of a presumed male performance advantage should be the default position. It thus follows that athletes with these XY DSDs hold male performance advantages. Since many authors of the commentary have acknowledged that male performance advantages result from developmental androgenization [3], that androgenization is a feature of certain XY DSDs [7], and that male androgenization justifies ineligibility in a protected female category [3], Williams et al.'s rejection of our proposal as unjustified on scientific grounds is contradictory. Williams et al. inappropriately “straw man” our position to criticize an assumed screening of minors. Our proposal does not advocate for this, nor do we set a target age. Rather, we believe that eligibility screening should occur early enough in an athlete's career to protect their privacy and dignity and avoid the ethical failures of the past [8]. Furthermore, Williams et al. overlook the reality that sex verification procedures are already used in many sports but routinely applied in an ad hoc manner that lacks standardization and is targeted based on suspicion. Notably, World Aquatics have introduced a cohort-wide requirement for athletes to certify their chromosomal sex to meet international eligibility. We are not, then, proposing novelty, but arguing for a more ethical approach that improves fairness and equitable treatment of all athletes. Maintaining the status quo enables the problems we have already seen to persist and will continue to result in significant harm to athletes. We propose that atypical screen results prompt immediate referral to clinical specialists, who typically conduct extensive anatomical, genetic, and endocrinological tests within established medical workflows to secure a diagnosis [9]. As this “standard medical care” is beyond the remit of sports federations, it is clinical specialists who must address the ethical challenges of delivering “invasive” and “potentially humiliating” care. As a final point on ethics, also misleading is Williams et al.'s characterization of screening as a coercive offer. Were this true, it would rule out eligibility or doping tests of any kind. Williams et al. raise concerns that cohort-wide sex screening would be costly and impractical. However, technological advances mean a simple sex screen would be inexpensive, require minimal equipment and could be completed in under 60 min. Implementation could be stratified and phased appropriately to spread cost, as has already been done in anti-doping programs. As we noted [2], cohort-wide screening is supported by 82% of female athletes [8], and ultimately, sport organizations have a duty to respect the internationally recognized human rights of girls and women to equality and non-discrimination in sport on the basis of sex [10]. We look forward to constructive discourse between scientists, sports associations, and other key stakeholders on this topic, including proposals from Williams et al. for alternative approaches that protect the integrity of female sport. We believe that a broader screening process with follow-up examinations in rare cases is scientifically sound, ethically justifiable and operationally feasible. The authors have nothing to report. The authors would like to make a joint conflict of interest statement in which they declare the following: Several authors have received payment to provide expert testimony related to this topic. Several authors have received payment for their consultancy work with sports organizations and/or companies. Several authors have received travel and accommodation expenses for speaking engagements related to this topic. Several authors have spoken in the mainstream media on this topic. Three authors (E.N.H., C.D., and J.P.) are unpaid advisors to advocacy organizations. The authors have nothing to report.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".