Machine Learning-Derived Predictors of Clinical Outcomes in CureGN Participants with FSGS
Bibliographic record
Abstract
Background: Focal segmental glomerulosclerosis (FSGS) is a common disease pattern with variable clinical presentations, outcomes, and etiologies, including immune-mediated, adaptive, and genetic. Much of the clinical heterogeneity remains unexplained. Methods: 722 participants with FSGS have been enrolled with 512 having whole genome sequencing (WGS) data, 276 with centralized biopsy scoring for conventional morphologic parameters, and 226 having both WGS and pathology scoring. Sequential ridge regression models were performed using: 1) baseline demographic, 2) social determinants of health, clinical, and genetic, 3) pathology morphologic features to predict time to kidney failure defined as two eGFR measurements <15mL/min, dialysis, or kidney transplantation. Ridge regression allows complex modelling while avoiding overfitting. Missing variables were imputed where possible. Discrimination was assessed using integrated area under the curve (iAUC) and variables were ranked by absolute value of standardized coefficients. Results: Among participants with WGS data, 92 had a high-risk APOL1 genotype and 32 had monogenic kidney disorders. Model discrimination improved with the stepwise addition of features from model 1 to 3 (iAUC = 0.86, 0.93, 0.95 respectively). Among the 31 most predictive features 14 were protective (1 demographic, 6 clinical, 7 morphologic) and 17 were risk factors (2 demographic, 7 clinical, 1 genetic, 7 morphologic). The top ranked protective factors included high eGFR, high serum albumin, low rates of interstitial fibrosis and tubular atrophy (IFTA), high calcium, diffuse mesangial hypercellularity, high hemoglobin and sodium, tip variant FSGS, and no inflammation in areas of IFTA. Top ranked risk factors included high levels of IFTA, high proteinuria, tubular microcystic changes, hypertension, collapsing variant FSGS, inflammation in areas of IFTA, high levels of segmental and globally sclerosed glomeruli, thrombotic diagnosis, high potassium, edema, high-risk APOL1 genotype, private health insurance, high levels of arteriosclerosis, black race, use of RAAS inhibitors, and young age at onset. Conclusion: Machine learning methodologies that integrate a broad range of clinical, genetic, and pathology data allow for the identification of parameters that are predictive of clinical outcome in FSGS. Funding: NIDDK Support
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.009 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".