Prognostic stratification in early‐stage hepatocellular carcinoma: Imaging biomarkers are needed
Bibliographic record
Abstract
Early-stage hepatocellular carcinoma (HCC), defined as a single tumour (stage 0 if ≤2 cm; stage A if >2 cm) or ≤3 tumours ≤3 cm (stage A) as per the Barcelona Clinic Liver Cancer (BCLC) classification,1 is eligible for curative treatment including liver transplantation, surgical resection, and percutaneous ablation. Although percutaneous ablation has been shown to offer overall survival rates comparable to those of surgical resection for small lesions, the latter remains the standard therapeutic option for larger lesions (>3 cm) provided that the patient is a good surgical candidate.2, 3 When both surgical resection and percutaneous ablation are feasible, the choice between the two possibilities mainly relies on liver function, portal pressure, age, comorbidities, and local expertise and, except for alpha-fetoprotein level, no factors related to the tumour aggressiveness are routinely taken into account. For obvious reasons, the presence of microvascular invasion cannot be currently assessed in tumours treated by ablation. However, tumour differentiation is accessible by percutaneous biopsy, which should reinforce its systematic performance. It is therefore urgent to develop pretherapeutic biomarkers to capture the heterogeneity of early-stage HCC, risk stratify their prognosis to personalize the treatment and identify those that could potentially benefit from adjuvant treatment. In this issue of Liver International, Wang et al. aimed to develop a state-of-the-art deep learning model to predict microvascular invasion (MVI) using pre-treatment magnetic resonance imaging (MRI) in patients with solitary tumours ≤3 cm.4 The originality of this work was to use a cohort of patients with surgically resected HCC to train the deep learning model to predict the presence of MVI before testing the model in a cohort of patients with ablated HCC with recurrence-free survival and overall survival as the primary outcomes. With a very large dataset of 696 patients and a limited imbalance between positive (28.5%) and negative cases of MVI to train the model, the latter achieved high performance in the validation cohort of surgically resected HCC (AUC of .901 and .816 in BCLC A and 0 HCC, respectively). Interestingly, when tested in the ablation cohort, the imbalance between cases at high risk of MVI and those at low risk was similar (30.6%) with no significant difference in size between HCCs with high risk of MVI and those without. In the ablation cohort, the recurrence-free survival rates of patients with high MVI risk were 57.1% at 1 year, 30.7% at 2 years, 13.1% at 3 years, and 2.6% at 5 years, which were significantly lower than those of patients predicted without MVI (87.8% at 1 year, 80.4% at 2 years, 71.3% at 3 years, and 56.0% at 5 years, p < .001). The 1-, 3-, and 5-year overall survival rates were 90.9%, 68.2%, and 49.1% for patients with high MVI risk, which were also significantly lower than 98.4%, 92.2%, and 81.5% for patients predicted without MVI, respectively (p < .001). Using a stepwise multivariate Cox regression analysis, this MVI biomarker was shown to be an independent risk factor for a lower recurrence-free survival rate, in addition to alpha-fetoprotein >20 ng/mL and unfavourable tumour location. The latter is most likely explained by technical difficulties and the heat-sink effect as it was the only independent risk factor for local tumour progression. Interestingly, the MVI biomarker was significantly associated with intrahepatic distance recurrence (32.3% vs. 9.6% at 1 year, 57.3% vs. 13.0% at 2 years, 71.0% vs. 20.1% at 3 years, 76.3% vs. 32.8% at 5 years), which could reinforce its clinical relevance as a companion biomarker of adjuvant therapies. If no significant risk factor was found for extrahepatic metastasis, this can be explained by the extremely limited number of positive cases (5/180). The overall C-index of the multivariate Cox regression model for evaluating recurrence-free survival was .73. In addition to the large size of the datasets and the robustness of the deep learning methodology, another strength reinforcing the generalisability of the model is the relative heterogeneity of the included MR images. Although all MR scanners in the training centre were from the same vendor, the acquisition parameters of the different MRI sequences were heterogeneous enough to ensure the generalisability of the model to another dataset of MR images from a different vendor. Furthermore, this study reinforces the importance of exploiting the full potential of MRI to capture all the tumour specificities. In the same way, radiologists analyse HCCs from all MRI sequences, the deep learning model developed in this study performed better when integrating a multiphase approach than a single-phase approach (AUC of .883 vs. .685–.763 in the validation cohort). Clinical transferability is the main future challenge of this study, which can also be stated for most of the artificial intelligence studies. To be implemented in the clinical routine, the computation and use of imaging-based biomarkers must be as simple as possible. It is unrealistic to think that, with the increasing number of imaging studies, radiologists will have the time to make manual annotations on all the MRI phases. Imaging-based biomarkers should be just a click away. This requires certain technical issues to be resolved, such as the misregistration of sequences mainly due to inconsistent breathing, or automated segmentation of the tumour. It should also be remembered that one objective of developing this type of biomarker is to provide prognosis stratification for the accurate selection of patients with early-stage HCC who will most benefit from adjuvant therapies, therefore avoiding side effects if no benefit is expected, to achieve the best clinical outcome/cost ratio. Therefore, although relevant from a computational perspective, the best clinical threshold of stratification biomarkers may not be found by maximizing sensitivity and specificity as it is done in this study. Finally, a classic limitation to the applicability of such deep learning models to populations with different epidemiology is the high prevalence of patients with chronic hepatitis B virus (93.0%–96.1%) and the absence of cirrhosis in almost half of the patients in the training cohort. It can be noted that a major clinical difference between the ablation cohort and the surgical training cohort was the higher prevalence of cirrhosis in the ablation cohort (72.2% vs. 44.8%). Nevertheless, it is usual to find such differences as there is a systematic selection bias in patients treated with ablation (older age, higher total bilirubin, lower albumin, etc.). To return to everyday clinical challenges, the stratification of patients' prognosis at the initial diagnosis of their tumour remains an unmet need. For instance, the deep learning model proposed here, based on imaging features of tumour aggressiveness, has the potential to assist physicians in choosing the most appropriate treatment based on the predicted MVI risk. Indeed, surgery may be a more appropriate treatment for small HCCs (<3 cm) in surgical candidates than ablation if MVI is likely to be present around the tumour. Additionally, for nonsurgical candidates or unresectable tumours, detecting features of MVI on pre-therapeutic imaging could lead to more aggressive interventional radiology treatments (such as multipolar ablation or combined ablation + transarterial chemoembolization) to ensure wider ablation margins and reduce the risk of loco-regional recurrence (Figure 1). However, both curative-intent treatments are associated with a high risk of intrahepatic or distant recurrence, reported up to 70% overall at 5 years.2, 5 In this context, multiple clinical trials have been conducted to evaluate the impact of local (i.e., transarterial chemoembolization) and systemic (i.e., sorafenib or more recently immunotherapies) adjuvant therapies on progression-free survival and overall survival.6 Recently, an interim analysis of the randomized phase III clinical trial (IMbrave050) of adjuvant atezolizumab + bevacizumab for patients at high risk of recurrence following resection or ablation has demonstrated a significant improvement in recurrence-free survival.7 The high risk of recurrence was well defined in a surgical cohort based on the size, number of resected tumours, the poor differentiation of the tumour and the presence of micro- or macrovascular invasion on the surgical specimen. Interestingly, in the subgroup of patients treated by ablation, the benefit of the adjuvant treatment was less clear, and the criteria to define the high risk of recurrence were different, relying solely on tumour size and number. Pending the results of other ongoing phase II/III trials, no international recommendation currently endorses the use of (neo)adjuvant treatment in the setting of percutaneous ablation especially for single small tumours which is the specific population studied in the Wang et al. study. Hence, these approaches based on imaging biomarkers are relevant and can help identify a target population for future studies where adjuvant therapies would be tested in the context of percutaneous ablation. In conclusion, this study is a promising step forward in developing prognosis stratification biomarkers that could be used to optimize the curative treatment approach and to identify patients with early-stage HCC that might benefit from adjuvant therapies after tumour ablation. Prospective clinical trials are needed to refine and test this imaging-based biomarker before it can be implemented in clinical practice. The authors do not have any disclosures to report. Data sharing is not applicable to this article as no datasets were generated or analysed during the current study.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".