Improving risk stratification and detection of early HCC using ultrasound-based deep learning models
Bibliographic record
Abstract
Background & Aims: Hepatocellular carcinoma (HCC) surveillance programs are suboptimal. We aimed to design an ultrasound-based deep learning model for HCC risk stratification (STARHE-RISK) and early-stage HCC detection (STARHE-DETECT) in patients with compensated advanced chronic liver disease (cACLD). Methods: This prospective multicentric study included 403 adult patients with cACLD of all causes enrolled in a surveillance program for at least 6 months without prior history of HCC. STARHE-RISK was trained on ultrasound cine clips of the non-tumoral liver parenchyma using two classes: cases (n = 152 patients with early-stage HCC; 137/152 [82%] male; median age 63 years) and controls (n = 170 patients without HCC at inclusion and during a subsequent 1-year follow-up; 120/170 [71%] male; median age 69 years). STARHE-DETECT was trained on tumour ultrasound cine clips. The training/validation and testing sets were stratified according to potential confounders, and 50 patients who were balanced in both groups were allocated to the independent testing set based on sample size calculation. Statistical analysis included classification and detection metrics. Results: = 0.004]) for predicting a patient at high risk of HCC development. STARHE-DETECT achieved a 0.67 mAP10, a 0.68 sensitivity (95% CI 0.47-0.85), and a 0.82 specificity (95% CI 0.69-0.91) for detecting early-stage HCC. Conclusions: STARHE-RISK and STARHE-DETECT achieved robust performances for HCC risk stratification and early-stage HCC detection, respectively. They could become valuable surveillance tools and pave the way for a risk-based personalised surveillance program. Impact and implications: STARHE-RISK is a reliable ultrasound-based deep learning model for hepatocellular carcinoma (HCC) risk stratification in patients with compensated advanced chronic liver disease and can be associated with complementary scores integrating clinical and blood parameters. STARHE-DETECT could become a complementary tool to visual assessment for radiologists and sonographers in HCC surveillance. Both models are based on simple and easy-to-perform ultrasound cine clip acquisitions. This study paves the way for a risk-based personalised surveillance program that will not ultimately rely on a single test but rather on a combination of approaches mixing clinical, biological, and radiological data. Clinical Trials Registration: The study is registered at ClinicalTrials.gov (NCT04802954).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".