Predicting low testosterone in aging men: a systematic review
Bibliographic record
Abstract
BACKGROUND: Physicians diagnose and treat suspected hypogonadism in older men by extrapolating from the defined clinical entity of hypogonadism found in younger men. We conducted a systematic review to estimate the accuracy of clinical symptoms and signs for predicting low testosterone among aging men. METHODS: We searched the MEDLINE and Embase databases (January 1966 to July 2014) for studies that compared clinical features with a measurement of serum testosterone in men. Three of the authors independently reviewed articles for inclusion, assessed quality and extracted data. RESULTS: Among 6053 articles identified, 40 met the inclusion criteria. The prevalence of low testosterone ranged between 2% and 77%. Threshold testosterone levels used for reference standards also varied substantially. The summary likelihood ratio associated with decreased libido was 1.6 (95% confidence interval [CI] 1.3-1.9), and the likelihood ratio for absence of this finding was 0.72 (95% CI 0.58-0.85). The likelihood ratio associated with the presence of erectile dysfunction was 1.5 (95% CI 1.3-1.8) and with absence of erectile dysfunction was 0.83 (95% CI 0.76-0.91). Of the multiple-item instruments, the ANDROTEST showed both the most favourable positive likelihood ratio (range 1.9-2.2) and the most favourable negative likelihood ratio (range 0.37-0.49). INTERPRETATION: We found weak correlation between signs, symptoms and testosterone levels, uncertainty about what threshold testosterone levels should be considered low for aging men and wide variation in estimated prevalence of the condition. It is therefore difficult to extrapolate the method of diagnosing pathologic hypogonadism in younger men to clinical decisions regarding age-related testosterone decline in aging men.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.016 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.003 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".