Characterization of Focal Liver Masses: A Multicenter Comparison of Contrast‐Enhanced Ultrasound, Computed Tomography, and Magnetic Resonance Imaging
Bibliographic record
Abstract
OBJECTIVES: To demonstrate the usefulness of contrast-enhanced ultrasound (CEUS) for the evaluation of focal liver masses via a direct comparison to standard ultrasound and computed tomography/magnetic resonance imaging (CT/MRI). METHODS: A cohort of 214 patients with previously undiagnosed focal liver masses were included from 5 different centers. Each patient was imaged using CEUS and CT and/or MRI. Anonymized and randomized images were interpreted by 4 separate blind readers from 3 of the participating centers (2 readers for CEUS and 2 readers for CT/MRI). Readers were blinded to patient demographics and past medical history. Readers were asked to decide if the lesion was benign or malignant, provide a final diagnosis for the lesion, and provide a confidence interval. Results were compared to truth standard from pathology or expert consensus. RESULTS: In determination of malignancy, CEUS had a sensitivity of 95%, specificity of 82%, PPV of 82%, NPV of 95%, statistically better than standard ultrasound (sensitivity 82%, specificity 56%, PPV 60%, NPV 78%) with P < .01 and not statistically different from CT (sensitivity 90%, specificity 73% PPV 81%, NPV 86%) or MRI (sensitivity 85%, specificity 79%, PPV 68%, NPV 91%) with P ≥ .01. In assigning a final diagnosis, CEUS had an accuracy of 78% statistically better than standard ultrasound (46%) with P < .01 and not statistically different from CT (68%) or MRI (71%) with P > .01. CONCLUSIONS: In the evaluation of focal liver lesions, both for determination of malignancy and in accuracy of final diagnosis, CEUS performs better than standard ultrasound and at least equivalent to both CT and MRI.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".