The Improved Kidney Risk Score in ANCA-Associated Vasculitis for Clinical Practice and Trials
Bibliographic record
Abstract
SIGNIFICANCE STATEMENT: Reliable prediction tools are needed to personalize treatment in ANCA-associated GN. More than 1500 patients were collated in an international longitudinal study to revise the ANCA kidney risk score. The score showed satisfactory performance, mimicking the original study (Harrell's C=0.779). In the development cohort of 959 patients, no additional parameters aiding the tool were detected, but replacing the GFR with creatinine identified an additional cutoff. The parameter interstitial fibrosis and tubular atrophy was modified to allow wider access, risk points were reweighted, and a fourth risk group was created, improving predictive ability (C=0.831). In the validation, the new model performed similarly well with excellent calibration and discrimination ( n =480, C=0.821). The revised score optimizes prognostication for clinical practice and trials. BACKGROUND: Reliable prediction tools are needed to personalize treatment in ANCA-associated GN. A retrospective international longitudinal cohort was collated to revise the ANCA renal risk score. METHODS: The primary end point was ESKD with patients censored at last follow-up. Cox proportional hazards were used to reweight risk factors. Kaplan-Meier curves, Harrell's C statistic, receiver operating characteristics, and calibration plots were used to assess model performance. RESULTS: Of 1591 patients, 1439 were included in the final analyses, 2:1 randomly allocated per center to development and validation cohorts (52% male, median age 64 years). In the development cohort ( n =959), the ANCA renal risk score was validated and calibrated, and parameters were reinvestigated modifying interstitial fibrosis and tubular atrophy allowing semiquantitative reporting. An additional cutoff for kidney function (K) was identified, and serum creatinine replaced GFR (K0: <250 µ mol/L=0, K1: 250-450 µ mol/L=4, K2: >450 µ mol/L=11 points). The risk points for the percentage of normal glomeruli (N) and interstitial fibrosis and tubular atrophy (T) were reweighted (N0: >25%=0, N1: 10%-25%=4, N2: <10%=7, T0: none/mild or <25%=0, T1: ≥ mild-moderate or ≥25%=3 points), and four risk groups created: low (0-4 points), moderate (5-11), high (12-18), and very high (21). Discrimination was C=0.831, and the 3-year kidney survival was 96%, 79%, 54%, and 19%, respectively. The revised score performed similarly well in the validation cohort with excellent calibration and discrimination ( n =480, C=0.821). CONCLUSIONS: The updated score optimizes clinicopathologic prognostication for clinical practice and trials.
Stored with the screening record, where it is evidence for the labels above.
How this classification was reachedexpand
The three-model screen
all 5,600 screened works →All three models called this out of scope.
Development and validation of a clinical kidney risk prediction score in ANCA-associated vasculitis; a clinical prognostic tool, not a study of research practice.
The study develops and validates a clinical risk score for ANCA-associated vasculitis.
Clinical prognostic score for ANCA-associated kidney disease; object is disease risk, not trial methodology.
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.079 | 0.246 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.003 | 0.005 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.004 | 0.002 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.007 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".