External validation of the Toronto hepatocellular carcinoma risk index in a Swedish population
Bibliographic record
Abstract
•The THRI is a simple and non-invasive method to estimate 5- and 10-year HCC risk.•This was the largest validation of the THRI to date.•The THRI had a modest discriminative ability and was well-calibrated.•However, the THRI could only identify few patients at low risk of HCC, limiting its clinical use. Background & AimsThe Toronto hepatocellular carcinoma (HCC) risk index (THRI) is a predictive model to determine the risk of HCC in patients with cirrhosis. This study aimed to externally validate the THRI in a Swedish setting to investigate whether it could identify patients not requiring HCC surveillance.MethodsFrom 2004-2017, 2,491 patients with cirrhosis at the Karolinska University Hospital were evaluated. Patients were classified into low-, intermediate- and high-risk groups for future HCC according to the THRI. Harrell’s C-index, calibration-in-the-large, calibration slope and goodness-of-fit estimates were calculated to assess model discrimination and calibration. Cox proportional hazards regression was used to determine the risk of HCC.ResultsMost patients were male (n = 1,638, 66%). The most common etiologies of cirrhosis were steatohepatitis (n = 1,182, 48%) followed by viral hepatitis (n = 987, 40%). In all, 131 patients (5.3%) were designated as low risk for HCC. Harrell’s C-index was 0.69. Calibration-in-the-large (0.11), calibration slope (1.24, not different from 1, p = 0.66) and goodness-of-fit showed good model calibration. Patients in the high-risk group had a 7.1-fold (95% CI 2.9–17.2) higher risk of HCC and patients in the intermediate-risk group had a 2.5-fold (95% CI 1.0–6.3) higher risk compared to the low-risk group.ConclusionsIn a Swedish setting, the THRI could differentiate between low- and high-risk of HCC development. However, because the low-risk group was relatively small (5.3%), the clinical applicability of the THRI could be limited.Lay summaryThe Toronto hepatocellular carcinoma (HCC) risk index (THRI) is a novel prediction model used to stratify patients with cirrhosis based on future risk of HCC. In this study, the THRI was validated in an external cohort using the TRIPOD guidance. Few patients were identified as low-risk, and the THRI had a modest discriminative ability, limiting its clinical applicability. The Toronto hepatocellular carcinoma (HCC) risk index (THRI) is a predictive model to determine the risk of HCC in patients with cirrhosis. This study aimed to externally validate the THRI in a Swedish setting to investigate whether it could identify patients not requiring HCC surveillance. From 2004-2017, 2,491 patients with cirrhosis at the Karolinska University Hospital were evaluated. Patients were classified into low-, intermediate- and high-risk groups for future HCC according to the THRI. Harrell’s C-index, calibration-in-the-large, calibration slope and goodness-of-fit estimates were calculated to assess model discrimination and calibration. Cox proportional hazards regression was used to determine the risk of HCC. Most patients were male (n = 1,638, 66%). The most common etiologies of cirrhosis were steatohepatitis (n = 1,182, 48%) followed by viral hepatitis (n = 987, 40%). In all, 131 patients (5.3%) were designated as low risk for HCC. Harrell’s C-index was 0.69. Calibration-in-the-large (0.11), calibration slope (1.24, not different from 1, p = 0.66) and goodness-of-fit showed good model calibration. Patients in the high-risk group had a 7.1-fold (95% CI 2.9–17.2) higher risk of HCC and patients in the intermediate-risk group had a 2.5-fold (95% CI 1.0–6.3) higher risk compared to the low-risk group. In a Swedish setting, the THRI could differentiate between low- and high-risk of HCC development. However, because the low-risk group was relatively small (5.3%), the clinical applicability of the THRI could be limited.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".