Predicting treatment outcome in congenital adrenal hyperplasia using urine steroidomics and machine learning
Bibliographic record
Abstract
OBJECTIVE: Treatment monitoring of individuals with congenital adrenal hyperplasia (CAH) remains unsatisfactory. Comprehensive 24 h urine steroid profiling provides detailed insight into adrenal steroid pathways. We investigated whether 24 h urine steroid profiling can predict treatment control in children and adolescents with CAH using machine learning (ML). DESIGN: Prospective observational cohort study. METHODS: This study included children with 21-hydroxylase deficiency. On 24 h urines of 2 consecutive visits 40 steroids were measured by gas chromatography-mass spectrometry. Treatment outcome was clinically classified as undertreated, optimally treated or overtreated. We used sparse partial least squares discriminant analysis (sPLS-DA) to investigate prediction of treatment outcome. We computed area under the ROC-curve (AUC) of 2 sPLS-DA models: (1) using only 24 h urine metabolites and (2) adding clinical variables. RESULTS: We included 112 visits (68 optimal, 44 undertreatment) from 59 patients: 27 (46%) girls, 46 (78%) classic CAH, and 19 (32%) prepubertal. Mean age at first visit was 11.9 ± 4.0 years and mean BMI SDS 0.6 ± 1.1. SPLS-DA using 24 h urine metabolites showed clear clustering of optimally treated patients on 2 components, while undertreated patients were more heterogeneous (AUC 0.88). The model selected pregnanetriol and 17α-hydroxypregnanolone contributing to excluding optimal treatment and 5 metabolites contributing to excluding undertreatment: 17β-estradiol, cortisone, tetrahydroaldosterone, androstenetriol, and etiocholanolone. Addition of clinical variables marginally improved classification (AUC 0.90). CONCLUSIONS: Using ML on 24 h urine steroid profiling predicted treatment outcome in children with CAH, even in the absence of clinical data, suggesting that routine comprehensive 24 h urine steroid profiling could improve treatment monitoring in CAH.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".