Notice bibliographique
Résumé
CITATION Yoo D, Goutaudier V, Divard G, et al. An automated histological classification system for precision diagnostics of kidney allografts. Nat Med. 2023;29(5):1211-1220. https://doi.org/10.1038/s41591-023-02323-6 Diagnosis of allograft rejection remains one of the most challenging issues in the clinic. The Banff classification was initially developed as an international consensus more than 30 years ago for diagnosis of allograft rejection based primarily on histologic features of graft biopsies. Since allograft rejection is not an absence or presence of disease but rather a continuously evolving process, consensus thresholds must be established first for rendering a diagnosis of rejection while avoiding overdiagnosis or underdiagnosis to guide treatment decisions. The Banff classification nowadays represents complex consensus rules derived from associations between histopathology lesions, allograft function and outcome, donor-specific antibody detection, and certain biomarkers (C4d and gene expression). Over time, the multimodal Banff lesions and rules have increased in complexity to a point where humans have become unable or unwilling to follow those complex rules, leading to considerable interpractitioner and intrapractitioner variabilities in rendering and interpreting Banff diagnosis. Such variabilities have become a significant issue that has potentially deleterious impact on patient care. In a recent article in Nature Medicine, Yoo et al from the Paris Institute for Transplantation and Organ Regeneration report a computer-based automation approach toward applying the complex Banff rules in allograft rejection diagnostics. The authors assembled a large consortium involving multiple transplant centers in Europe and North America and developed a computerized decision-support system. They translated all Banff classification rules and potential diagnostic scenarios into a computer algorithm that can automatically assign diagnoses to kidney allograft biopsies. In other words, they developed a computerized system faithfully following the most current 2019 Banff rules. The authors tested this decision-support system, which is not an artificial intelligence system but a sophisticated “if-then” algorithm, for reclassifying rejection-related diagnoses for adult and pediatric kidney transplant recipients in 3 international multicenter cohorts and 2 large prospective clinical trials, which included 4409 biopsies from 3054 patients followed in 20 transplant referral centers in Europe and North America. They found that, in the adult kidney transplant population, the Banff automation system reclassified 83 out of 279 (29.75%) antibody-mediated rejection cases and 57 out of 105 (54.29%) T cell–mediated rejection cases into other diagnostic categories, whereas 237 out of 3,239 (7.32%) biopsies diagnosed as nonrejection by pathologists were reclassified as rejection. On the other hand, 7.3% of adults with no rejection diagnosis were reclassified using the Banff automation system into various types of rejection diagnoses. In the pediatric population, the reclassification rates into other diagnostic categories were 8 out of 26 (30.77%) for antibody-mediated rejection and 12 out of 39 (30.77%) for T cell–mediated rejection (Fig.). Clearly, a substantial fraction of diagnoses rendered by pathologists were reclassified by the Banff automation system. The authors pointed out that the main causes for misclassifications by pathologists included (1) misinterpretation of the Banff classification (28.8%); (2) complex rejection cases with misapplication of the Banff diagnostic rules (48.3%); and (3) changes in classification rules over time (22.9%), among which 16.7% were related to the fact that pathologists used an outdated version of the classification at the time of assessing biopsies. But most importantly, the authors found that reclassification of the initial diagnoses rendered by pathologists applying the Banff automation system was associated with improved risk stratification of long-term allograft outcomes. Specifically, patients diagnosed by a pathologist as nonrejection but reclassified using the Banff automation system as rejection displayed worse graft survival than patients without rejection. Moreover, patients diagnosed by a pathologist as rejection but reclassified by the Banff automation system as nonrejection showed excellent graft survival, similar to patients without rejection diagnosed by both pathologists using the Banff automation system. It is laudable that the authors made the full code and data openly available to others to reproduce their Banff automation system at https://www.synapse.org/BanffAutomationSystem. They also deployed their Banff automation system online to offer potential users a free and user-friendly application to verify the adequacy of their diagnoses using the 2019 Banff rules (https://transplant-prediction-system. shinyapps.io/Banff_automation/). Clearly, this study calls attention to the potential benefits of computerized decision-support tools in routine patient care. By “simply” following current Banff rules using an automated Banff classification process, they demonstrate improved transplant patient care by correcting human errors and standardizing allograft rejection diagnoses. This is exciting since the system utilizes input variables from the Banff lesion scores generated locally by pathologists as well as clinical and laboratory variables produced by local transplant centers, all known to be prone to significant interobserver and intraobserver and laboratory variabilities. One might speculate what can be achieved if those variabilities are further reduced or even neutralized by advanced digital technologies and artificial intelligence tools for Banff lesions, especially when combined with molecular diagnostics and prognostication systems like the ibox, where all variables are integrated into multidimensional electronic health records. Nevertheless, the algorithm developed by Yoo et al from the Paris Institute for Transplantation and Organ Regeneration is a major milestone toward data-driven and evidence-based practice in an increasingly complex health care environment that, today, likely exceeds most humans’ intellectual ability to process all relevant information accurately .
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,013 | 0,050 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,001 | 0,002 |
| Communication savante | 0,003 | 0,003 |
| Science ouverte | 0,002 | 0,001 |
| Intégrité de la recherche | 0,007 | 0,009 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,005 | 0,003 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».