Phenotypes and Prognostic Subgroups Derived by the RejectClass Clustering Algorithm Are Not Fully Reproducible in an Independent Multicenter Study
Bibliographic record
Abstract
Kidney transplant biopsies are crucial in clinical care of kidney transplant patients. Pathology of kidney transplant biopsies is evaluated histologically according to the international Banff consensus classification.1,2 The Banff schema defines kidney transplant rejection diagnosis and its subtypes according to ordinal scores of 16 histological or immunohistochemical lesions seen in biopsy tissues (glomerulitis, peritubular capillaritis, interstitial inflammation in unscarred cortical parenchyma, tubulitis in unscarred cortical parenchyma, arteritis, total cortical interstitial inflammation, interstitial inflammation in scarred cortical parenchyma, tubulitis in scarred cortical parenchyma, transplant glomerulopathy, mesangial matrix increase, interstitial fibrosis, tubular atrophy, arterial fibrous intimal thickening, arteriolar hyalinosis, complement 4d immunostaining [C4d], simian virus 40 immunostaining-polyomavirus replication/load level). As both acute and chronic features are evaluated, diagnosis in pathology reports often includes multiple diagnostic categories that are not mutually exclusive plus activity and chronicity status of the rejection state reflected by the rejection type and subtype (eg, acute T cell–mediated rejection [TCMR] plus chronic active antibody-mediated rejection [AMR]). Although the Banff system is very useful to enable diagnosis of rejection, its major disadvantages include complexity of its rules and difficulties in the application of multiple lesion scores, which in turn result in low reproducibility of Banff lesion scores/diagnostic categories as well as complex pathology results with many overlapping or mixed entities. To overcome the difficulties of the Banff system and better assist clinical treatment-decisions, a recent study by Vaulet et al3 identified data-driven novel phenotypes of acute kidney transplant rejection using semisupervised K-means clustering (RejectClass acute clusters) in 2 retrospective patient cohorts. The researchers identified 6 clinically relevant clusters of biopsies using only acute Banff lesion scores and donor-specific antibody (DSA) data at time of biopsy, which are weighted by survival data. These data-driven clusters were trained in 3510 biopsies from 936 patients (Leuven cohort) and validated in an external cohort of 3835 biopsies from 1989 patients (Lyon and Paris cohorts) retrospectively. The identified 6 clusters were prognostically different and appeared to simplify the Banff diagnostic categories. Each patient was placed into 1 cluster eliminating borderline or suspicious for rejection or overlapping phenotypes. The same group later generated chronic allograft clusters using chronic Banff lesions, which were weighted by their time-dependent association with graft failure in a large retrospective training and validation cohorts.4 This yielded 4 chronic clusters that are associated with graft outcomes. In addition to acute or chronic clusters, the authors also came up with a quantitative sum inflammation severity or chronicity score calculated per biopsy. The authors concluded that their clustering approach improves survival predictions of the Banff categories and offers a quantitative evaluation of rejection types and their activity and chronicity severity to assist clinicians in decision-making, especially when Banff labels are not definitive. Furthermore, they made the RejectClass Clustering algorithm available on a website to be tested by the international community. In this issue of Transplantation, a study by von Samson-Himmelstjerna et al5 tests the validity of the RejectClass clusters in an independent multicenter retrospective cohort of 616 indication biopsies from 441 patients followed in 4 German transplant centers. The evaluation of biopsies in this independent cohort by the previously established RejectClass algorithm showed only a partially comparable phenotype content in clusters when compared with the original RejectClass studies. Although some clusters were associated with increased graft loss, this study does not fully reproduce all phenotypes and their prognostic stratification. Some of the important points include as follows: The original RejectClass study contained 6 acute clusters: cluster 1, no inflammation/no rejection; cluster 2, DSA-negative glomerulitis; cluster 3, TCMR-like; cluster 4, DSA-positive no/mild inflammation; cluster 5, AMR-like; and cluster 6, mixed AMR- and TCMR-like. Compared with cluster 1, all other clusters were associated with poor graft survival in the original study. In the current study, cluster 2 included DSA-negative microvascular inflammation biopsies with higher peritubular capillaritis scores, more positive C4d, and higher tubulitis and interstitial inflammation scores compared with the original study; thus, this cluster appears as DSA-negative AMR/mixed rejection diagnosed by glomerulitis/peritubular capillaritis lesions and minimal focal C4d (Banff labels were AMR or mixed AMR-TCMR in 9 of 11 cluster 2 biopsies). Of note, only 27% of cluster 2 cases were C4d score 1 and the rest were 4d negative; thus, it is unclear how they diagnosed Banff AMR in 9 of 11 biopsies in cluster 2 (I would speculate likely based on historical DSA detection before the time of biopsy). Thus, the original interpretation of cluster 2 as a novel phenotype which is not recognized by Banff (DSA-negative glomerulitis) in the RejectClass study does not seem to be true in this new study. Instead, cluster 2 seems to include cases of DSA-negative AMR captured by microvascular inflammation and minimal C4d staining. Borderline category is eliminated in both the current and the original studies and placed into cluster 1 (no rejection), cluster 3 (TCMR-like) or cluster 4 (DSA-positive no-mild inflammation). Intriguingly, AMR biopsies were spread into all 6 acute clusters. The reason for this spread is unclear, but perhaps related with heterogeneity of inflammatory or chronicity lesion scores in the AMR cases (eg, AMR with low inflammation in cluster 1 or 2). This pattern was not seen in the original study where AMR biopsies were spread into mostly cluster 4 and 5. Thus, this observation seems like a significant difference between the studies. As compared with the original study, only cluster 3 and 6 (TCMR-like and mixed rejection-like) were associated with reduced graft survival in the current study. Clusters 2, 4, and 5 were not associated with reduced graft survival. Although these latter clusters had poor prognosis in the original study, this lack of validation may be explained by smaller sample size and lower statistical power in the current study. The original study showed that clustering of biopsies improved prediction of graft survival when compared with the Banff classification. This point was not validated in the current study. As all Banff rejection categories (TCMR, AMR, borderline, and mixed) were associated with reduced graft survival, but it was not the case with all RejectClass clusters. Both the original study and the new study showed that chronic clusters of interstitial fibrosis and tubular atrophy (chronic clusters 2–3) and transplant glomerulopathy (chronic cluster 4) were associated with reduced graft survival. This is not a new finding as these histological lesions are well-known to be associated with poor kidney survival. Both the quantitative inflammation and chronicity sum scores derived from the clustering were associated with graft survival in the current study as in the original studies. Again, this is not a surprising finding, as most of the Banff histological lesions are associated with graft prognosis. Neither the original studies nor the current study included nonrejection diseases into analysis or clusters. This is a drawback of these studies as some common diseases such as recurrent glomerulonephritis or BK virus nephritis are associated with reduced graft survival. In summary, RejectClass clustering is a useful method to stratify a given biopsy into one of the few clusters to predict inflammatory and chronicity burden and associated prognostic subgroup. However, it remains unclear whether this quantitative method improves diagnostic classification and prognostication. The current study in this issue of Transplantation does not show that RejectClass is superior to Banff for survival stratification. Elimination of the borderline category by the RejectClass clusters seems intriguing. However, similar to the Banff, RejectClass clusters are also heterogeneous and carry ambiguity. It remains unclear if clusters are composed of biologically distinct diseases or simply represent stratifying patients into different prognostic subgroups based on a mixture of lesions that carry poor prognostic significance. In the acute clusters study, chronic lesion scores were ignored and not entered into the clustering algorithm and similarly, acute lesion scores were ignored in the chronic cluster study. This does not make sense as both acute and chronic lesions concurrently present in many biopsies and thus both must be considered for diagnosis, activity/staging and prognostication purposes. We hope that future clustering studies will be able to achieve this goal in a reproducible manner.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.045 | 0.080 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.003 |
| Bibliometrics | 0.002 | 0.003 |
| Science and technology studies | 0.003 | 0.001 |
| Scholarly communication | 0.005 | 0.002 |
| Open science | 0.003 | 0.004 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.006 | 0.003 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".