MétaCan
Menu
Back to cohort
Record W1608537416 · doi:10.1002/hep.26936

Screening for Liver Cancer: Another Piece of the Puzzle?

2013· letter· en· W1608537416 on OpenAlexaffabout
Morris Sherman

Bibliographic record

VenueHepatology · 2013
Typeletter
Languageen
FieldMedicine
TopicLiver Disease Diagnosis and Treatment
Canadian institutionsToronto General HospitalUniversity of Toronto
Fundersnot available
KeywordsMedicineLiver cancerCancerInternal medicine

Abstract

fetched live from OpenAlex

Despite the fact that most hepatologists accept the need for screening for hepatocellular carcinoma (HCC), it remains a controversial issue.1, 2 As with cancer screening for so many other tumors, proving that HCC screening is effective has been difficult. Ideally, a cancer screening program, to be effective, should result in a decrease in disease-specific mortality (i.e., fewer people should die of cancer). This can only be demonstrated with certainty in a prospective, randomized, controlled trial (RCT) in which one group is screened and the other not screened, with mortality as an outcome. For HCC, there are two such trials, both conducted in China.3, 4 The first failed to show a mortality difference. However, this study used only alpha-fetoprotein (AFP) as the screening test. Subsequent knowledge has clearly shown that AFP is an inadequate screening test. Furthermore, the study report did not indicate what treatments were applied to patients in whom HCC was detected; but in the era when this study was conducted, China was a poor country, with the costs of surgery being borne by the patient. This likely led to treatment not being applied to all those who might have benefited. The second study did show a substantial reduction in mortality, but this study has been criticized on the basis that the analysis was statistically incorrect. This criticism does decrease confidence in the results of this study, but does not say that screening is ineffective. There is considerable evidence, albeit of lower quality, that supports that screening is likely to reduce HCC mortality. This uncertainty has implications for public policy. Payers want certainty. To date, with the exception of Japan and South Korea, no governmental agency has recommended HCC screening. Should payers wait for randomized, controlled data, the only data that bring certainty, before recommending and paying for HCC screening? For many reasons, although desirable, initiation of an RCT of HCC screening is unlikely, and if initiated, completion is also unlikely. Because there is at least suggestive evidence of a substantial benefit to HCC screening, payers and government agencies have to consider the issue and make their decisions based on available evidence. What evidence, short of RCT data, might be sufficient to convince payers and government agencies that HCC screening should be recommended? Case-control studies showing earlier diagnosis (stage migration) are not proof of efficacy nor are retrospective studies, many of which are subject to lead-time bias. Prospective, nonrandomized studies with mortality as an endpoint are better than retrospective studies, but are all subject to bias. This bias can be, to some extent, mitigated by using propensity score matching (i.e., matching study subject and control on criteria that are important to the risk of developing HCC). This principle has been well demonstrated in a study of the efficacy of entecavir in reducing HCC incidence.5 In this issue of Hepatology, a group from Taiwan report on a study in which a risk score is applied to identify a cohort of subjects who are deemed, by virtue of having an elevated risk score, to be at higher risk of HCC.6 HCC screening with a one-time ultrasound (US) was offered. The comparison groups were a simultaneous cohort being the other inhabitants of the same counties who were not invited to participate, as well as a historical control group. The study showed that those who underwent screening, albeit one time, had a lower HCC mortality than those who did not, despite the fact that the screened group was at higher risk for HCC and therefore at higher risk of HCC-related death. Several aspects of this study merit comment. The Taiwan study selected patients on the basis of two risk scores: one for patients with viral hepatitis and one for those free of viral hepatitis. The viral hepatitis risk score included AFP, alanine aminotransferase, and platelet count <150 ×109/mL, whereas the other score included the presence or absence of diabetes instead of viral markers. Both scores included terms for age and gender. However, above a cutoff that led to the invitation to screen, no further stratification was done. The analysis of mortality compared the “invited group” apparently whether or not they accepted the invitation to undergo screening with the general population and with the historical controls who also did not undergo screening. Approximately 21% refused screening. This is analogous to an intention-to-treat analysis and would tend to bias the results against screening being effective. Although the study results are interesting, they probably do not contribute much to the debate about the efficacy of HCC screening. First, subjects underwent only a single US examination, rather than sequential imaging over time, which would be current practice. Second, the use of the risk score is problematic because it has not been externally validated. However, the risk score seemed to have served its purpose in that incidence of HCC and prevalence of cirrhosis was higher in those with high risk scores and correlated with the risk score. Finally, the follow-up was short, only 15 months. A longer follow-up is probably required to confirm reduction in morality, because patients with screen-detected HCC would be expected to live longer than 15 months whether or not screening has any effect on mortality, simply because of earlier detection (lead-time bias). Inclusion of transaminases in the risk score would identify a population with liver disease without specifying that cirrhosis is present. Inclusion of a low platelet count in the score will enrich the population for cirrhosis (i.e., for a population at risk for HCC). As such, the risk score might be useful if it were to be applied to a general population to identify silent cirrhosis not necessarily for screening, but for more liver-directed medical care. However, to do so, we need to know how sensitive the score is—in other words, what proportion of all at-risk subjects does the score not identify. We are not given this information. There are better validated HCC risk scores, including one from Taiwan. The REACH B score was derived from the REVEAL study and validated externally in China, Hong Kong, and Korea.7, 8 There are two other risk scores for hepatitis B (Chinese University9 and the GAG-HCC score10), both derived in Asian cohorts, but not externally validated. There are no equivalent scores for patients with hepatitis B who might be infected with genotypes other than B and C. There are no equivalent scores for patients who do not have viral hepatitis. A validated risk score can be used to assess an individual patient's HCC risk and therefore to determine whether HCC screening for that individual is appropriate or not. However, the risk score used in this study cannot be used in this manner. This is because, as described in the article, it does not indicate the incidence of HCC in the population identified by the risk score. This information is required to determine whether screening was cost-efficient and what cut-off score would generate a risk sufficient to warrant screening. The risk score includes a low platelet count and so probably identifies some patients with established cirrhosis and some degree of portal hypertension—who would be candidates for screening even without a risk score. Thus, overall, although this article is interesting, the results from this study are not generalizable to most populations that might be candidates for HCC screening. Although it shows a benefit to HCC screening, it uses a target population and a once-off methodology that are unlike current guidelines. It also raises another point, namely, that those who had a “low risk” and were not invited to undergo screening may still have had a risk sufficiently high to warrant screening using other criteria to guide the decision. It is also unfortunate that the causes of mortality were not given. We cannot assume that all deaths were the result of HCC. Given the age of the population, it is likely that there were competing causes of death. The question of whether HCC screening will decrease HCC mortality therefore remains uncertain, and this study does not resolve the question of whether HCC screening is effective or not. Morris Sherman, M.B., B.Ch., Ph.D., F.R.C.P(C). University of Toronto Toronto General Hospital Toronto, Ontario, Canada

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesInsufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.148
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0010.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.051
GPT teacher head0.288
Teacher spread0.237 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designNot applicable
Domainnot available
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations2
Published2013
Admission routes2
Has abstractyes

Explore more

Same venueHepatologySame topicLiver Disease Diagnosis and TreatmentFrench-language works237,207