Predicting obstruction risk using common ultrasonography parameters in paediatric hydronephrosis with machine learning
Bibliographic record
Abstract
OBJECTIVE: To sensitively predict the risk of renal obstruction on diuretic renography using routine reported ultrasonography (US) findings, coupled with machine learning approaches, and determine safe criteria for deferral of diuretic renography. PATIENTS AND METHODS: Patients from two institutions with isolated hydronephrosis who underwent a diuretic renogram within 3 months following renal US were included. Age, sex, and routinely reported US findings (laterality, kidney length, anteroposterior diameter, Society for Fetal Urology [SFU] grade) were abstracted. The drainage half-times were collected from renography and stratified as low risk (<20 min, primary outcome), intermediate risk (20-60 min), and high risk of obstruction (>60 min). A random Forest model was trained to classify obstruction risk, here named the 'Artificial intelligence Evaluation of Renogram Obstruction' (AERO). Model performance was determined by measuring area under the receiver-operating-characteristic curve (AUROC) and decision curve analysis. RESULTS: A total of 304 patients met the inclusion criteria, with a median (interquartile range) age of diuretic renogram at 4 (2-7) months. Of all patients, 48 (16%) were low risk, 102 (33%) were intermediate risk, 156 (51%) were high risk of obstruction based on diuretic renogram. The AERO achieved a binary AUROC of 0.84, multi-class AUROC of 0.74 that was superior to the SFU grade, and external validation (n = 64) binary AUROC of 0.76. The most important features for prediction included age, anteroposterior diameter, and SFU grade. We deployed our application in an easy-to-use application (https://sickkidsurology.shinyapps.io/AERO/). At a threshold probability of 30%, the AERO would allow 66 more patients per 1000 to safely avoid a renogram without missing significant obstruction compared to a strategy in which a renogram is routinely performed for SFU Grade ≥3. CONCLUSIONS: Coupled with machine learning, routine US findings can improve the criteria to determine in which children with isolated hydronephrosis a diuretic renogram can be safely avoided. Further optimisation and validation are required prior to implementation into clinical practice.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".