Re: Indeterminate 1-2 Cm Nodules Found on Hepatocellular Carcinoma Surveillance: Biopsy for All, Some, or None?
Bibliographic record
Abstract
We thank Drs. Iavarone and Sangiovanni for the interest in our study,1 but caution the authors with regard to several points in their letter. First, their conclusions are based on a small sample size of 36 indeterminate nodules. While they calculate sensitivity and specificity of 44% and 55%, respectively, using our proposed criteria, the 95% confidence interval was not reported. We calculate their confidence interval to be 21%-69% for sensitivity and 32%-76% for specificity. Small sample sizes lead to real uncertainty. Second, the authors report that 16/35(46%) of their malignant nodules did not demonstrate a typical enhancement pattern on imaging. This rate is well above those for <2 cm nodules reported by Forner et al. (15%),2 Leoni et al. (7%),3 or us (25%).1 The substantial lower sensitivity reported by Iavarone and Sangiovanni may be a result of their small sample size, but should lead to reexamination of their methodology. It is unclear whether the authors used the critical delayed phase in assessment of washout.4 They also used fixed imaging times after contrast injection for all magnetic resonance imaging (MRI) and an indeterminate number of computed tomography (CT) scans, compromising phase timing. Third, we disagree with fine-needle biopsy (FNB) as the reference standard. The substantial false-negative rate of biopsy is ignored by the authors in both their study and their letter. FNB, as opposed to core biopsy, further compromises the diagnosis of very early hepatocellular carcinomas (HCCs) due to its inability to detect architectural changes such as sinusoidal invasion.5 Biopsy relies on the judgment of a pathologist to predict future behavior of a nodule, whereas close imaging follow-up demonstrates actual behavior: growth. For the purposes of a study, long-term stability represents a stronger reference standard than biopsy. Finally, Iavarone and Sangiovanni worry that our proposed criteria may lead to "significantly delayed" diagnosis in a proportion of patients. Our role is detection and treatment of only malignancies that cause morbidity or shorten life. If the term "significant" is to be used outside its statistical definition, it should be within such a framework. In the setting of a competing potentially fatal disease (cirrhosis), the treatment of "very early HCCs" has yet to be justified. We invite the authors and others to perform a new prospective trial to independently evaluate our proposals. Korosh Khalili MD*, Morris Sherman MD , * Department of Medical Imaging, University of Toronto, Toronto, ON, Canada, Department of Gastroenterology, University of Toronto, Toronto, ON, Canada.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.017 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.002 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.005 | 0.003 |
| Insufficient payload (model declined to judge) | 0.008 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".