Biomarker implementation: Evaluation of the decision‐making impact of CXCL10 testing in a pediatric cohort
Bibliographic record
Abstract
Abstract Background Children are at high risk for subclinical rejection, and kidney biopsy is currently used for surveillance. Our objective was to test how novel rejection biomarkers such as urinary CXCL10 may influence clinical decision‐making to indicate need for a biopsy. Methods A minimum dataset for standard decision‐making to indicate a biopsy was established by an expert panel and used to design clinical vignettes for use in a survey. Pediatric nephrologists were recruited to review the vignettes and A) estimate rejection risk and B) decide whether to biopsy; first without and then with urinary CXCL10/Cr level. Accuracy of biopsy decisions was then tested against the biopsy results. IRA was assessed by Fleiss Kappa (κ) for binary choice and ICC for probabilities. Results Eleven pediatric nephrologists reviewed 15 vignettes each. ICC of probability assessment for rejection improved from poor (0.28, P < .01) to fair (0.48, P < .01) with addition of CXCL10/Cr data. It did not, however, improve the IRA for decision to biopsy (K = 0.48 and K = 0.43, for the comparison). Change in clinician estimated probability of rejection with additional CXCL10/Cr data was correlated with CXCL10/Cr level (r 2 = 0.7756, P < .0001). Decision accuracy went from 8/15 (53.3%) cases to 11/15 (73.3%) with CXCL10/Cr, although improvement did not achieve statistical significance. Using CXCL10/Cr alone would have been accurate in 12/15 cases (80%). Conclusion There is high variability in decision‐making on biopsy indication. Urinary CXCL10/Cr improves probability estimates for risk of rejection. Training may be needed to assist nephrologists in better integrate biomarker information into clinical decision‐making.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".