Predicting Liver Transplant Patient Outcomes. Is a Validated Model Enough?
Notice bibliographique
Résumé
The choice of a good candidate for liver transplantation (LT) is based on the individual risk benefit ratio. As organ shortage affects LT programs and the number of patients in need of a graft always exceeds the number of donors, the access to LT depends also on the severity and urgency of other transplant candidates. To optimize the selection process and give priority for LT, the concepts of urgency, utility, and transplant benefit have been raised in the past years.1 The model for end-stage liver disease (MELD) score is used in the majority of allocation systems as it highly correlates with the mortality on the waiting list (urgency-based system). Interestingly, it has also been associated with transplant benefit.2 However, MELD score does not always capture the real severity of the patient, and MELD exceptions have been integrated in allocation models.3 These models are constantly changing to improve selection policies raising the concern that better models that account for more factors are needed.4 Molinari5 in this issue of Transplantation validated the liver transplant risk score (LTRS) that was published by the same group few years ago.6 External validation of a score is truly needed, and the authors executed this with the maximum scientific rigor. The score, based on 5 pre-LT variables (recipient’s age, MELD score, body mass index, history of diabetes, and need for dialysis), predicts 90-days post-LT survival (utility based system) and aims to improve the selection process. Moreover, in the recent manuscript, the authors found that the model also predicts 1-year mortality and 4-year survival. The LTRS is calculated when the patient is referred for LT and allows the stratification of patients in 5 risk categories (90-d mortality score 0 = 2.7% versus ≥5 = 9.3% and 1-y mortality score 0 = 5.5% versus ≥5 = 15.4%). The validation on a large independent population testified the solidity of the score. Although the LTRS is robust, some drawbacks need to be highlighted. As for the MELD score, the LTRS is based only on objectives variables, which is valuable as it avoids subjective evaluation, but it also limits the score. It does not take into account parameters that contribute to describe the severity of the liver disease. Ascites, encephalopathy, frailty, hypoalbuminemia, hyponatremia, female gender, hepatopulmonary syndrome are only some of the variables, which can severely impact LT candidates.6 A machine-learning algorithm was used to select variables with the best correlation with 90-day mortality. Undoubtedly, this is a great instrument, objective and precise, that can help to work with big data and will likely be present in the transplant field in the years to come. However, a machine-learning model applied only to pre-LT variables at a given point seems too simplistic to analyze the complexity of a LT candidate and predict post-LT outcome. Combination of donor, recipient, and transplant factors to predict early post-LT mortality has been analyzed previously using machine-learning methodology.7 Furthermore, intention-to-treat analysis is crucial when analyzing the outcomes of LT recipients. Indeed, the lack of data on waitlist mortality in the current study and the absence of a competing-risk analysis limit the applicability of their score. Specific scores have been used over time to define criteria for inclusion of hepatocellular carcinoma (HCC) patients in LT waiting lists. Patients with HCC are in constant competition with patients without HCC, and to facilitate their access to LT, different models based on exception points have been used.8 The work of Molinari et al included both patients with and without malignancies; however, the pretransplant severity, 1-year mortality, and in particular, 4-year survival for HCC and non-HCC patients are difficult to predict with a same model. This model captures the patient’s risk at a determinate time point, at the time of LT evaluation or at the time of the listing. It is well known that the equity and urgency of a single individual compared with the other patients on the waiting list change over time from the moment of the inscription on the waiting list to the moment of the allocation of a specific organ. An interesting example is given by patients with acute-on-chronic liver failure grade 3 at listing that can improve to a lower grade of acute-on-chronic liver failure at transplant with a significantly higher post-LT survival.9 The evolution of a patient’s severity seems in need of a dynamic score, and the validity of the LTRS at different time points while the patient is on the waiting list has not yet been tested. The power of the LTRS in predicting post-LT mortality and survival was not compared with the existent scores, and it is difficult to know whether the performance is better than the MELD score alone for example. Interestingly, the patients in the worse category had a fairly good outcome with a 4-year survival rate of 75%. It would be difficult, therefore, to discriminate a patient with a high LRTS. Another limitation is that the validation was carried out on the United Network for Organ Sharing database as for the development of the model, so futile transplant cannot be identified. This disputes the utility of the score in the clinical practice. Finally, the applicability of this score in jurisdictions different than the United States is also in doubt. As observed in this interesting initiative, although we are looking forward to mathematical scores as objective tools to guide us in the everyday difficult decision-making process, scores often present several limitations. Molinari et al need to be congratulated for their study. Future studies are needed to evaluate dynamic models and to take advantage of machine-learning methodology to account for waitlist changes. The final aim to improve the outcome prediction after LT and consequently, to ameliorate LT organ allocation remains a challenge.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,023 | 0,086 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,000 |
| Méta-épidémiologie (sens large) | 0,002 | 0,001 |
| Bibliométrie | 0,001 | 0,001 |
| Études des sciences et des technologies | 0,000 | 0,001 |
| Communication savante | 0,005 | 0,003 |
| Science ouverte | 0,002 | 0,001 |
| Intégrité de la recherche | 0,002 | 0,004 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,004 | 0,002 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».