Bibliographic record
Abstract
CITATION Yoo D, Goutaudier V, Divard G, et al. An automated histological classification system for precision diagnostics of kidney allografts. Nat Med. 2023;29(5):1211-1220. https://doi.org/10.1038/s41591-023-02323-6 Diagnosis of allograft rejection remains one of the most challenging issues in the clinic. The Banff classification was initially developed as an international consensus more than 30 years ago for diagnosis of allograft rejection based primarily on histologic features of graft biopsies. Since allograft rejection is not an absence or presence of disease but rather a continuously evolving process, consensus thresholds must be established first for rendering a diagnosis of rejection while avoiding overdiagnosis or underdiagnosis to guide treatment decisions. The Banff classification nowadays represents complex consensus rules derived from associations between histopathology lesions, allograft function and outcome, donor-specific antibody detection, and certain biomarkers (C4d and gene expression). Over time, the multimodal Banff lesions and rules have increased in complexity to a point where humans have become unable or unwilling to follow those complex rules, leading to considerable interpractitioner and intrapractitioner variabilities in rendering and interpreting Banff diagnosis. Such variabilities have become a significant issue that has potentially deleterious impact on patient care. In a recent article in Nature Medicine, Yoo et al from the Paris Institute for Transplantation and Organ Regeneration report a computer-based automation approach toward applying the complex Banff rules in allograft rejection diagnostics. The authors assembled a large consortium involving multiple transplant centers in Europe and North America and developed a computerized decision-support system. They translated all Banff classification rules and potential diagnostic scenarios into a computer algorithm that can automatically assign diagnoses to kidney allograft biopsies. In other words, they developed a computerized system faithfully following the most current 2019 Banff rules. The authors tested this decision-support system, which is not an artificial intelligence system but a sophisticated “if-then” algorithm, for reclassifying rejection-related diagnoses for adult and pediatric kidney transplant recipients in 3 international multicenter cohorts and 2 large prospective clinical trials, which included 4409 biopsies from 3054 patients followed in 20 transplant referral centers in Europe and North America. They found that, in the adult kidney transplant population, the Banff automation system reclassified 83 out of 279 (29.75%) antibody-mediated rejection cases and 57 out of 105 (54.29%) T cell–mediated rejection cases into other diagnostic categories, whereas 237 out of 3,239 (7.32%) biopsies diagnosed as nonrejection by pathologists were reclassified as rejection. On the other hand, 7.3% of adults with no rejection diagnosis were reclassified using the Banff automation system into various types of rejection diagnoses. In the pediatric population, the reclassification rates into other diagnostic categories were 8 out of 26 (30.77%) for antibody-mediated rejection and 12 out of 39 (30.77%) for T cell–mediated rejection (Fig.). Clearly, a substantial fraction of diagnoses rendered by pathologists were reclassified by the Banff automation system. The authors pointed out that the main causes for misclassifications by pathologists included (1) misinterpretation of the Banff classification (28.8%); (2) complex rejection cases with misapplication of the Banff diagnostic rules (48.3%); and (3) changes in classification rules over time (22.9%), among which 16.7% were related to the fact that pathologists used an outdated version of the classification at the time of assessing biopsies. But most importantly, the authors found that reclassification of the initial diagnoses rendered by pathologists applying the Banff automation system was associated with improved risk stratification of long-term allograft outcomes. Specifically, patients diagnosed by a pathologist as nonrejection but reclassified using the Banff automation system as rejection displayed worse graft survival than patients without rejection. Moreover, patients diagnosed by a pathologist as rejection but reclassified by the Banff automation system as nonrejection showed excellent graft survival, similar to patients without rejection diagnosed by both pathologists using the Banff automation system. It is laudable that the authors made the full code and data openly available to others to reproduce their Banff automation system at https://www.synapse.org/BanffAutomationSystem. They also deployed their Banff automation system online to offer potential users a free and user-friendly application to verify the adequacy of their diagnoses using the 2019 Banff rules (https://transplant-prediction-system. shinyapps.io/Banff_automation/). Clearly, this study calls attention to the potential benefits of computerized decision-support tools in routine patient care. By “simply” following current Banff rules using an automated Banff classification process, they demonstrate improved transplant patient care by correcting human errors and standardizing allograft rejection diagnoses. This is exciting since the system utilizes input variables from the Banff lesion scores generated locally by pathologists as well as clinical and laboratory variables produced by local transplant centers, all known to be prone to significant interobserver and intraobserver and laboratory variabilities. One might speculate what can be achieved if those variabilities are further reduced or even neutralized by advanced digital technologies and artificial intelligence tools for Banff lesions, especially when combined with molecular diagnostics and prognostication systems like the ibox, where all variables are integrated into multidimensional electronic health records. Nevertheless, the algorithm developed by Yoo et al from the Paris Institute for Transplantation and Organ Regeneration is a major milestone toward data-driven and evidence-based practice in an increasingly complex health care environment that, today, likely exceeds most humans’ intellectual ability to process all relevant information accurately .
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.050 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.003 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.007 | 0.009 |
| Insufficient payload (model declined to judge) | 0.005 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".