Optimizing US for HCC surveillance
Bibliographic record
Abstract
Ultrasound is the primary imaging modality used for surveillance of patients at risk for HCC. In 2017, the American College of Radiology Liver Imaging Reporting and Data Systems (ACR LI-RADS) introduced US LI-RADS to standardize the performance, interpretation, and reporting of US for HCC surveillance, with the algorithm recently updated as LI-RADS US Surveillance v2024. The American Association for the Study of Liver Diseases (AASLD) recommends reporting both the examination-level LI-RADS US Category as well as the US Visualization Score. The US Category conveys the overall findings of the exam and primarily determines follow up recommendations. The US Visualization Score conveys the expected sensitivity of the test and stratifies patients into appropriate surveillance pathways. One of the goals of routine surveillance is the detection of HCC at an early, potentially curable stage. Therefore, optimizing US technique is of critical importance. Increasing North American and worldwide utilization of LI-RADS US Surveillance, which includes technical recommendations, through education and outreach will undoubtedly benefit patients undergoing US HCC surveillance.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.003 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".