A287 LEVERAGING MACHINE LEARNING TO IMPROVE THE DIAGNOSTIC ACCURACY OF ULTRASOUND SCREENING FOR HEPATOCELLULAR CARCINOMA
Bibliographic record
Abstract
Abstract Background Ultrasound screening stands out as the gold standard for hepatocellular carcinoma (HCC) detection, attributed to its broad accessibility, patient-friendly, and cost-efficient nature. Nevertheless, the five-year survival rate for HCC currently rests at 32.7%, with suboptimal screening being a key contributor to this. The recent years have witnessed the rise of machine learning models, powered by artificial neural networks. Among these, convolutional neural networks (CNNs) have taken the lead in revolutionizing medical image analysis, offering unprecedented success in predictive tasks and giving us hope for a brighter future in the fight against HCC Aims To train and test a machine learning algorithm using pre-trained CNNs to improve early detection of HCC through ultrasound screening. Methods In this retrospective study, 1835 charts of patients with chronic liver disease were reviewed: 346 with histologically confirmed HCC and 1457 with ultrasounds without HCC. A diagnosis of HCC was confirmed pathologically on biopsy or surgical resection, and/or radiographically with a Liver Imaging Reporting and Data System (LI-RADS) score of five on CT and/or MRI. Patients with benign lesions were required to have at least two ultrasounds three years apart that confirmed benign characteristics. Cases were excluded if they had a prior history of treated HCC, post-transplant HCC, or HCC with Barcelona Clinic Liver Cancer (BCLC) Stage B and above. All ultrasound images were reviewed by experienced radiologists, and segmented as liver lesions (HCC versus benign) and surrounding liver. Results A total of 149 patients have been included to date, comprising 72 with benign lesions, 73 with HCC, and four with both benign and malignant lesions. 224 lesions have been segmented, consisting of 87 HCC and 137 benign lesions. Candidate networks are under development and evaluation for the classification of liver lesions. Imaging pre-processing was performed such that the liver region of interest (ROI) and the lesion ROI were standardized, with 2-channel greyscale images on two separate channels. The algorithm was constructed using a per-lesion analysis. Initial testing has achieved an area under the curve (AUC) of 73.3%, 95% CI [68.6, 77.0] with 2 repetitions of 10-fold cross-validation. Conclusions Enhancing ultrasound screening for HCC is imperative for improving patient care. Analysis of remaining cases is ongoing. Future studies will be essential, including prospective evaluation and external validation. Funding Agencies None
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.010 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".