Historical Data Are Not Relevant to the Diagnostic Performance of Ultrasound in Surveillance for Hepatocellular Carcinoma
Bibliographic record
Abstract
In the review of Tzartzeva et al’s1Tzartzeva K, et al. Gastroenterology;154:1706-1718.Google Scholar recent meta-analysis assessing diagnostic performance of surveillance tests for hepatocellular carcinoma (HCC), 2 fundamental considerations need to be considered from an imager’s perspective that substantially affect the conclusions reached by the authors. These considerations are not immediately apparent to non-imagers and thus commonly result in incorrect assumptions. Because I believe, along with the authors, that HCC surveillance affects the outcomes of patients positively in this common and fatal disease, I feel that it is important to outline these 2 matters, albeit belatedly. The first consideration is regarding the inclusion of studies that directly compared ultrasound examination with computed tomography scanning/magnetic resonance imaging in the same patient and that are not bona fide surveillance studies. Such studies falsely underestimate the effective sensitivity of ultrasound examination, which is bolstered by the ability to scan developing tumors at multiple surveillance points. The median tumor volume doubling time of HCC has been calculated and is relatively long, being 76.8 days for patients with hepatitis B, 137.2 days for patients with hepatitis C, and 99.8 days for patients with nonviral hepatitis.2An C, et al. Clin Mol Hepatol;21:279-286.Google Scholar The same study has calculated that for a single tumor to grow from 1 to 5 cm in diameter, the upper limit of “early HCC,” it takes a median of 678.9 days. That finding is entirely consistent with clinical practice and in fact a recent abstract from Tzartzeva et al’s group reports a substantially longer median doubling time of 225 days (7.5 months).3Rich N.E. et al.Hepatology. 2018; 68: 530A-531AGoogle Scholar The significance of this long doubling time means most HCCs grow over years without surpassing the curable or “early” designation. This allows for multiple opportunities for ultrasound examinations to detect the tumor if the tumor is missed by the initial surveillance scan. This multiple application of ultrasound scans increases its effective sensitivity. Studies in which ultrasound examination is directly compared with computed tomography scanning and/or magnetic resonance imaging in the same patient will always underestimate the former’s effective sensitivity as some tumors are picked up earlier by the other imaging. Tzaretzeva’s analysis includes Kim et al’s 2016 study4Kim SY, et al. JAMA Oncol;3:456-463.Google Scholar in which the sensitivity of ultrasound for “early” HCC is reported as 25.6%. The inclusion of this single study, a clear outlier (for the stated reason), significantly alters the analysis of sensitivity by decade, a crucial point explained below. The second consideration is the rapidly changing ultrasound technology. A sonographer would never include ultrasound scans performed in the 1980s and 1990s in an analysis when the purpose is to affect practice today. The simple reason is that the technology and resulting quality of imaging are exponentially better now. Alpha fetoprotein (AFP) has not changed over the past 40 years, but ultrasound technology has. With the clinical introduction of harmonic and compound imaging in mid 2000s, there was a substantial jump in the performance of ultrasound imaging. In a multivariate analysis of determinants of tumor detection by ultrasound examination, we have previously shown that the odds of detection of very early HCC (<2 cm, Barcelona Clinic Liver Cancer stage 0) significantly increased after the installation of scanners with these technologies as compared with prior generation of top quality scanners (odds ratio, 2.4; 95% confidence interval, 1.1–5.3).5Khalili K, et al. Can J Gastroenterol Hepatol;29:267-273.Google Scholar This assertion is also clearly visible in the reported sensitivity of ultrasound for “early” HCC by Tzaretzeva et al. When excluding Kim et al from pool of studies (listed in their Figure 1B ), the early detection rate is 29.7% in studies published in the 1990s, 49.0% in 2000s, and 62.2% in 2010s (Figure 1). Furthermore, the increase in sensitivity from the 2000s to the 2010s is statistically significant (P = .03). The conclusions reached by Tzaretzeva et al of 47% sensitivity of ultrasound examination alone for the detection of early HCC is a gross underestimation owing to the inclusion of a substantial proportion of outdated studies. The sensitivity of studies published over the past 8 years demonstrate a pooled sensitivity of 62.2% for ultrasound examination alone versus 71.4% for examination ultrasound with AFP. Although in their selected studies AFP seemed to significantly improve the performance of ultrasound examination in the detection of early HCC (P = .004), the absolute improvement of 9.2% is much closer to the 6% reported by our group, using a cutoff of 20 ng/mL.5Khalili K, et al. Can J Gastroenterol Hepatol;29:267-273.Google Scholar The high false-positive rate of AFP limits the cost effectiveness of this surveillance strategy6Andersson K.L. et al.Clin Gastroenterol Hepatol. 2008; 6: 1418-1424Abstract Full Text Full Text PDF PubMed Scopus (156) Google Scholar and the incremental gain needs to be assessed at each center subject to diagnostic performance of ultrasound examination alone. I thank Dr. Morris Sherman in providing editorial advice. Surveillance Imaging and Alpha Fetoprotein for Early Detection of Hepatocellular Carcinoma in Patients With Cirrhosis: A Meta-analysisGastroenterologyVol. 154Issue 6PreviewSociety guidelines differ in their recommendations for surveillance to detect early-stage hepatocellular carcinoma (HCC) in patients with cirrhosis. We compared the performance of surveillance imaging, with or without alpha fetoprotein (AFP), for early detection of HCC in patients with cirrhosis. Full-Text PDF ReplyGastroenterologyVol. 157Issue 3PreviewOur recent systematic review evaluating ultrasound examination and alpha fetoprotein (AFP) for early detection of hepatocellular carcinoma (HCC) had 2 primary conclusions: (1) abdominal ultrasound examination has suboptimal sensitivity for early detection of HCC when used alone and (2) using AFP with ultrasound examination can significantly increase sensitivity for early HCC detection, with minimal loss in specificity.1 The letter from Dr Khalili regarding our study raises interesting points that warrant further discussion. Full-Text PDF
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.065 | 0.241 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.004 | 0.008 |
| Bibliometrics | 0.005 | 0.007 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.003 | 0.005 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.002 | 0.004 |
| Insufficient payload (model declined to judge) | 0.007 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".