MétaCan
Menu
Back to cohort
Record W3043393415 · doi:10.1097/tp.0000000000003354

Predicting Liver Transplant Patient Outcomes. Is a Validated Model Enough?

2020· letter· en· W3043393415 on OpenAlexaff
Eléonora De Martin, Gonzalo Sapisochín

Bibliographic record

VenueTransplantation · 2020
Typeletter
Languageen
FieldMedicine
TopicLiver Disease and Transplantation
Canadian institutionsToronto General Hospital
Fundersnot available
KeywordsLiver transplantationMedicineIntensive care medicineInternal medicineTransplantation

Abstract

fetched live from OpenAlex

The choice of a good candidate for liver transplantation (LT) is based on the individual risk benefit ratio. As organ shortage affects LT programs and the number of patients in need of a graft always exceeds the number of donors, the access to LT depends also on the severity and urgency of other transplant candidates. To optimize the selection process and give priority for LT, the concepts of urgency, utility, and transplant benefit have been raised in the past years.1 The model for end-stage liver disease (MELD) score is used in the majority of allocation systems as it highly correlates with the mortality on the waiting list (urgency-based system). Interestingly, it has also been associated with transplant benefit.2 However, MELD score does not always capture the real severity of the patient, and MELD exceptions have been integrated in allocation models.3 These models are constantly changing to improve selection policies raising the concern that better models that account for more factors are needed.4 Molinari5 in this issue of Transplantation validated the liver transplant risk score (LTRS) that was published by the same group few years ago.6 External validation of a score is truly needed, and the authors executed this with the maximum scientific rigor. The score, based on 5 pre-LT variables (recipient’s age, MELD score, body mass index, history of diabetes, and need for dialysis), predicts 90-days post-LT survival (utility based system) and aims to improve the selection process. Moreover, in the recent manuscript, the authors found that the model also predicts 1-year mortality and 4-year survival. The LTRS is calculated when the patient is referred for LT and allows the stratification of patients in 5 risk categories (90-d mortality score 0 = 2.7% versus ≥5 = 9.3% and 1-y mortality score 0 = 5.5% versus ≥5 = 15.4%). The validation on a large independent population testified the solidity of the score. Although the LTRS is robust, some drawbacks need to be highlighted. As for the MELD score, the LTRS is based only on objectives variables, which is valuable as it avoids subjective evaluation, but it also limits the score. It does not take into account parameters that contribute to describe the severity of the liver disease. Ascites, encephalopathy, frailty, hypoalbuminemia, hyponatremia, female gender, hepatopulmonary syndrome are only some of the variables, which can severely impact LT candidates.6 A machine-learning algorithm was used to select variables with the best correlation with 90-day mortality. Undoubtedly, this is a great instrument, objective and precise, that can help to work with big data and will likely be present in the transplant field in the years to come. However, a machine-learning model applied only to pre-LT variables at a given point seems too simplistic to analyze the complexity of a LT candidate and predict post-LT outcome. Combination of donor, recipient, and transplant factors to predict early post-LT mortality has been analyzed previously using machine-learning methodology.7 Furthermore, intention-to-treat analysis is crucial when analyzing the outcomes of LT recipients. Indeed, the lack of data on waitlist mortality in the current study and the absence of a competing-risk analysis limit the applicability of their score. Specific scores have been used over time to define criteria for inclusion of hepatocellular carcinoma (HCC) patients in LT waiting lists. Patients with HCC are in constant competition with patients without HCC, and to facilitate their access to LT, different models based on exception points have been used.8 The work of Molinari et al included both patients with and without malignancies; however, the pretransplant severity, 1-year mortality, and in particular, 4-year survival for HCC and non-HCC patients are difficult to predict with a same model. This model captures the patient’s risk at a determinate time point, at the time of LT evaluation or at the time of the listing. It is well known that the equity and urgency of a single individual compared with the other patients on the waiting list change over time from the moment of the inscription on the waiting list to the moment of the allocation of a specific organ. An interesting example is given by patients with acute-on-chronic liver failure grade 3 at listing that can improve to a lower grade of acute-on-chronic liver failure at transplant with a significantly higher post-LT survival.9 The evolution of a patient’s severity seems in need of a dynamic score, and the validity of the LTRS at different time points while the patient is on the waiting list has not yet been tested. The power of the LTRS in predicting post-LT mortality and survival was not compared with the existent scores, and it is difficult to know whether the performance is better than the MELD score alone for example. Interestingly, the patients in the worse category had a fairly good outcome with a 4-year survival rate of 75%. It would be difficult, therefore, to discriminate a patient with a high LRTS. Another limitation is that the validation was carried out on the United Network for Organ Sharing database as for the development of the model, so futile transplant cannot be identified. This disputes the utility of the score in the clinical practice. Finally, the applicability of this score in jurisdictions different than the United States is also in doubt. As observed in this interesting initiative, although we are looking forward to mathematical scores as objective tools to guide us in the everyday difficult decision-making process, scores often present several limitations. Molinari et al need to be congratulated for their study. Future studies are needed to evaluate dynamic models and to take advantage of machine-learning methodology to account for waitlist changes. The final aim to improve the outcome prediction after LT and consequently, to ameliorate LT organ allocation remains a challenge.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.023
metaresearch head score (Gemma)0.086
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Commentary · Consensus signal: none
Teacher disagreement score0.023
Threshold uncertainty score0.123

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0230.086
Meta-epidemiology (narrow)0.0020.000
Meta-epidemiology (broad)0.0020.001
Bibliometrics0.0010.001
Science and technology studies0.0000.001
Scholarly communication0.0050.003
Open science0.0020.001
Research integrity0.0020.004
Insufficient payload (model declined to judge)0.0040.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.029
GPT teacher head0.248
Teacher spread0.219 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2020
Admission routes1
Has abstractyes

Explore more

Same venueTransplantationSame topicLiver Disease and TransplantationFrench-language works237,207