Fair and Safe Eligibility Criteria for Women's Sport: The Proposed Testing Regime Is Not Justified, Ethical, or Viable
Bibliographic record
Abstract
In an invited editorial, Tucker et al. [1] addressed the eligibility controversy regarding the Paris 2024 Olympic boxing competition. They cited Lundberg et al. [2] concerning the in/eligibility of transgender women for the female sports category and identified performance differences between males and females alongside studies involving testosterone suppression. Several authors of the present letter also co-authored Lundberg et al. [2] and stand by that paper, but declined co-authorship of Tucker et al.'s editorial [1] and present here, with additional collaborators, challenges to that editorial. First, the editorial [1] did not acknowledge the absence of high-quality scientific data regarding sport performance advantage of athletes with XY differences of sex development (DSDs). Furthermore, individuals with DSDs possess one or more of numerous rare genetic mutations [3], causing wide variability in pubertal physical development relevant to sport performance between and within different DSDs. Consequently, no primary evidence base currently exists demonstrating athletic performance advantage to justify testing and regulation of an entire population of competitors. Second, despite that lacuna, Tucker et al. [1] propose mandatory genetic testing in women's sports worldwide in “early”, sub-elite competition. Figure 2 in Lundberg et al. [2] includes high-performing teenage athletes. If consistent with their thesis, “early” must therefore include minors. Mandatory genetic testing was abandoned in 1999 due to concerns about validity, financial cost, and practicality, as well as trauma and stigmatization for many athletes [4]. All these concerns remain and are amplified by the vastly increased number of younger athletes tested under Tucker et al.'s [1] proposal. The editorial [1] gives the impression that such tests are straightforward, “individual consent, confidentiality, and dignity… simple cheek swab… standard medical care”, but these assurances ignore the enormous problems such a testing regime would generate. Individual voluntary consent for genetic testing is a core principle [5]. Under Tucker et al.'s [1] proposed mandatory genetic testing for sport eligibility rather than healthcare, young athletes would not be presented with a genuine choice. Consent is only a coercive offer: comply with the test or never participate in competitive women's or girls' sport, even at sub-elite level. Furthermore, ethically responsible genetic counseling ensures people understand the potential consequences of receiving genetic test results before consenting [6] and provides comprehensive professional follow-up [7]. Who would fund or produce this worldwide army of counseling expertise? For those undergoing follow-up clinical examination and genome sequencing (only providing a genetic diagnosis in ~50% of cases [8]), how would the devastation of young athletes' personal identity and self-esteem, and the alarm caused to their families [9], be managed? The resultant duty of care of these athletes will fall to the sport federations mandating such assessments, without any realistic prospect of being fulfilled. What, precisely, does follow-up clinical examination involve? Tucker et al. [1] say only “standard medical care”. They do not state transparently that it begins with clinical examination to assess clitoromegaly, symmetry of external genital structures, presence/absence of breast development, extent of sexual hair, involves palpation of genitalia, and so forth [9]. Indeed, as part of a “Level 1 Assessment”, World Athletics regulations [10] require such clinical examination by gynecologists or physicians following current guidance [9]. The concerns outlined above also apply to the World Athletics regulations, but even these affect only a small proportion of the athletes that would have to undergo such invasive and potentially humiliating examination under the editorial's [1] proposal. We agree with the editorial's [1] criticism of the current approach in some sports, which “invites targeted testing based on allegation, suspicion, subjective assessment, and bias” and broad discussion is required to develop more appropriate regulations. However, the proposed mandatory testing of all young women and girls in sport is not justified by scientific evidence, has limited ethical defensibility, and is not an operationally viable proposition. The authors would like to make a joint conflict of interest statement in which we declare the following: Several authors have received payment to provide expert testimony related to this topic and/or served as pro bono expert witnesses in relevant proceedings before the Court of Arbitration for Sport. Several authors have received payment for work with sports organisations. Several authors have received travel and accommodation expenses for speaking engagements related to this topic. Several authors have spoken in the mainstream media on this topic. Several authors have received research funding from relevant organisations including the International Olympic Committee, the World Anti-Doping Agency and the US Anti-Doping Agency. Data sharing is not applicable to this article as no new data were created or analyzed in this study.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".