A critical discussion of pediatric gender measures to clarify the utility and purpose of “measuring” gender
Bibliographic record
Abstract
Background Pediatric gender clinics and researchers commonly use scales to measure different dimensions of gender (e.g. identity, dysphoria, satisfaction). There has been little investigation into the relevance and consumer acceptability of these scales within contemporary understandings and experiences of gender.Aims This study aimed to comparatively review and evaluate measures of gender used with children and adolescents, to inform the use of gender measures in pediatric populations.Methods A narrative review of the literature was conducted to identify measures that are used to describe dimensions of gender within pediatric populations. The measures were evaluated for their inclusivity, validity, and utility.Results 19 measures were identified. Our results found that most pediatric gender measures are not inclusive of non-binary genders, and do not accommodate some understandings and expressions of gender. Many are based on outdated terminology and stereotyped expectations of gender expression, and some are potentially distressing for the young person completing the measure. Some gender measures, used in conjunction with self-identification and as an adjunct to clinical interviews, hold clinical utility for understanding gender. If a measure is deemed clinically helpful, it is vital that the purpose of the measure is explained to the young person, and they are supported through the administration of the measure.Discussion This review is a guide for choosing gender measures for clinical practice or research purposes. Specialist gender services and researchers should aim to provide an open, accepting, and affirmative approach; any gender measure should be chosen with consideration of its validity, and whether the measure adds value over and above self-identification and talking together about gender. There is a need for the development, and validation in pediatric populations, of measures that ensure the inclusivity of non-binary genders, language tailored to target ages and timepoints in gender transition, and updated, culturally appropriate language and examples.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".