Evaluation of delayed puberty: what diagnostic tests should be performed in the seemingly otherwise well adolescent?
Bibliographic record
Abstract
Delayed puberty (DP) is defined as the lack of pubertal development by an age that is 2-2.5 SDs beyond the population mean. Although it generally represents a normal variant in pubertal timing, concern that DP could be the initial presentation of a serious underlying disorder has led to a diagnostic approach that is variable and may include tests that are unnecessary and costly. In this review, we examine available literature regarding the recommended diagnostic tests and aetiologies identified during the evaluation of youth with DP. We view this literature through the prism of the seemingly otherwise well adolescent. To provide further clinical context, we also evaluate the clinical and laboratory data from patients seen with DP in our centre over a 2-year period. The literature and our data reveal wide variability in the number of tests performed and raise the question of whether tests, other than gonadotropins, obtained in the absence of signs or symptoms of an underlying disorder are routinely warranted. Together this information provides a pragmatic rationale for revisiting recommendations calling for broad testing during the initial diagnostic evaluation of an otherwise healthy adolescent with DP. We highlight the need for further research comparing the utility of broader screening with a more streamlined approach, such as limiting initial testing to gonadotropins and a bone age, which, while not diagnostic, is often useful for height prediction, followed by close clinical monitoring. If future research supports a more streamlined approach to DP, then much unnecessary testing could be eliminated.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.024 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.001 |
| Bibliometrics | 0.003 | 0.001 |
| Science and technology studies | 0.001 | 0.003 |
| Scholarly communication | 0.003 | 0.006 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.004 | 0.006 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".