Detection of all-cause advanced hepatic fibrosis using an ensemble machine learning framework
Bibliographic record
Abstract
The hallmark of liver injury is liver fibrosis. Ultimately, significant fibrosis deposition results in cirrhosis, whereby there are changes in liver architecture with nodule formation and ultimately, disturbances to liver function.1Garcia-Tsao G Friedman S Iredale J Pinzani M Now there are many (stages) where before there was one: in search of a pathophysiological classification of cirrhosis.Hepatology. 2010; 51: 1445-1449Crossref PubMed Scopus (376) Google Scholar Although the clinical features of liver decompensation are well described (ie, ascites, spider naevi, jaundice, signs of hepatic encephalopathy), patients who have early cirrhosis often have no clinical signs and might be entirely asymptomatic.2Muir AJ Understanding the complexities of cirrhosis.Clin Ther. 2015; 37: 1822-1836Summary Full Text Full Text PDF PubMed Scopus (30) Google Scholar The avoidance and prevention of liver fibrosis is one of the key objective for liver clinicians and is achieved through direct treatments such as antivirals, or through lifestyle modifications (eg, alcohol avoidance and weight loss).3Rinella ME Sanyal AJ Management of NAFLD: a stage-based approach.Nat Rev Gastroenterol Hepatol. 2016; 13: 196-205Crossref PubMed Scopus (204) Google Scholar Liver fibrosis has implications with respect to treatment initiation, follow-up, and prognosis. Greater liver scarring, without intervention, suggests a greater risk of cirrhosis and its complications long term (eg, end-stage liver failure and hepatocellular carcinoma). The challenges for clinicians include how to identify and capture populations at risk of liver disease and how to apply a test that accurately (and cheaply) identifies those individuals. Historically, liver biopsy has been regarded as the gold standard to detect liver fibrosis, but clearly, liver biopsy is not practical on a population-based level nor is it a suitable screening tool due to its invasiveness, costs, and its associated risks. Because of these reasons, there has been a proliferation of non-invasive tests of liver fibrosis over the past two decades.4Vilar-Gomez E Chalasani N Non-invasive assessment of non-alcoholic fatty liver disease: clinical prediction rules and blood-based biomarkers.J Hepatol. 2018; 68: 305-315Summary Full Text Full Text PDF PubMed Scopus (292) Google Scholar, 5Houot M Ngo Y Munteanu M Marque S Poynard T Systematic review with meta-analysis: direct comparisons of biomarkers for the diagnosis of fibrosis in chronic hepatitis C and B.Aliment Pharmacol Ther. 2016; 43: 16-29Crossref PubMed Scopus (67) Google Scholar These tests have included scores derived from routinely available blood tests and patient factors such as fibrosis-4 index (FIB-4), non-alcoholic fatty liver disease (NAFLD) fibrosis score (NFS), aspartate aminotransferase-platelet-ratio-index (APRI); panels of markers of fibrogenesis such as enhanced liver fibrosis and Fibrotest; and imaging-based tests such as transient elastography6Chang PE Goh GB Ngu JH Tan HK Tan CK Clinical applications, limitations and future role of transient elastography in the management of liver disease.World J Gastrointest Pharmacol Ther. 2016; 7: 91-106Crossref PubMed Google Scholar and magnetic resonance elastography. The imaging-based tests are accurate but expensive and not routinely available. Many, non-invasive tests seem unable to differentiate between minimal and advanced fibrosis. In short, these tests do not provide the clinician with the information they need to make informed decisions, particularly in regions where there is poor access to further tests. Screening tools are necessary in many clinical settings and it would be helpful if methods could be developed to minimise the risk of indeterminate results. It is for this reason that work in The Lancet Digital Health by Soren Sabet Sarvestany and colleagues is welcome.7Sarvestany SS Kwong JC Azhie A et al.Development and validation of an ensemble machine learning framework for detection of all-cause advanced hepatic fibrosis: a retrospective cohort study.Lancet Digit Health. 2022; 4: e188-e199Summary Full Text Full Text PDF Scopus (2) Google ScholarThe researchers trained six machine learning algorithms (MLAs) using clinical information and liver biopsies from two sites in Canada (Toronto Liver Clinic, Toronto, ON, and McGill University Health Centre, Montreal, QC). The score was developed from a training set of 1703 liver biopsies and then validated in 508 biopsies across a range of fibrosis stages. The team excluded patients with decompensated cirrhosis. The authors acknowledged the controversy around the length of liver biopsy and portal tract numbers on histological interpretation,8Rockey DC Caldwell SH Goodman ZD Nelson RC Smith AD Liver biopsy.Hepatology. 2009; 49: 1017-1044Crossref PubMed Scopus (1465) Google Scholar and how the local population fibrosis stage prevalence could influence the study findings.9Poynard T Halfon P Castera L et al.Standardization of ROC curve areas for diagnostic evaluation of liver fibrosis markers based on prevalences of fibrosis stages.Clin Chem. 2007; 53: 1615-1622Crossref PubMed Scopus (208) Google Scholar The best MLA was an ensemble algorithm (ENS) of support vector machine, random forest classifier, gradient booster classifier, logistic regression, and artificial neural network. The MLA was compared against APRI, FIB-4, NAFLD-fibrosis score NFS, and transient elastography.The MLA was superior to APRI, FIB-4, and NFS across the ranges of fibrosis (AUROC curves for F0 vs F4 and F01 vs F4 the ENS ranges were 0·827 [95 CI 0·751–0·895] to 0·803 [0·733–0·863] vs 0·820 [0·743–0·891] to 0·772 [0·708–0·836] for FIB-4; 0·764 [0·678–0·850] to 0·715 [0·639–0·794] for APRI; and 0·872 [0·802–0·931] to 0·833 [0·775–0·894] vs 0·822 [0·745–0·900] to 0·758 [0·682–0·835] for NFS). Yet, the ENS was not always better than transient elastography (0·773 [0·699–0·834] vs 0·826 [0·758–0·889]) but ENS performs better than NFS.MLAs have been developed before to accurately predict liver fibrosis stages.10Hashem S Esmat G Elakel W et al.Comparison of machine learning approaches for prediction of advanced liver fibrosis in chronic hepatitis C patients.IEEE/ACM Trans Comput Biol Bioinformatics. 2018; 15: 861-868Crossref PubMed Scopus (55) Google Scholar So, what does this new MLA offer to clinicians? Indeterminate scores are frustrating to patients and clinicians alike, leading to further investigations which could be unnecessary, harmful, and cause patient anxiety. So any test which reduces this uncertainty is welcome. This model could have a role in primary care and less well-developed health care settings. Moreover, this model could minimise referrals to secondary care services and provide earlier assurances to patients. But, to be useful, the tool must be easy to use and accessible. The development of online calculators and websites has been of benefit to clinicians, and this would be necessary if this tool were to be used widely. Furthermore, the result findings must be easy to interpret (eg, correlation with fibrosis stages) if the tool is to have clinical utility.In summary, the work by Sabet Sarvestany and colleagues provides insights into how an MLA can improve the accuracy of existing models of non-invasive liver fibrosis to further refine information, reduce the number of indeterminate scores, and allow better discrimination between those patients not at risk and those at risk of significant liver scarring and cirrhosis. Whether this test becomes universally used remains to be seen, but that will be determined by how easy it is to use and interpret, its accessibility, accuracy, and its use across a spectrum of liver conditions. But this work is of interest and relevance to the field. As such, it gives a hint of the future direction, not just of liver disease but medicine more broadly. The hallmark of liver injury is liver fibrosis. Ultimately, significant fibrosis deposition results in cirrhosis, whereby there are changes in liver architecture with nodule formation and ultimately, disturbances to liver function.1Garcia-Tsao G Friedman S Iredale J Pinzani M Now there are many (stages) where before there was one: in search of a pathophysiological classification of cirrhosis.Hepatology. 2010; 51: 1445-1449Crossref PubMed Scopus (376) Google Scholar Although the clinical features of liver decompensation are well described (ie, ascites, spider naevi, jaundice, signs of hepatic encephalopathy), patients who have early cirrhosis often have no clinical signs and might be entirely asymptomatic.2Muir AJ Understanding the complexities of cirrhosis.Clin Ther. 2015; 37: 1822-1836Summary Full Text Full Text PDF PubMed Scopus (30) Google Scholar The avoidance and prevention of liver fibrosis is one of the key objective for liver clinicians and is achieved through direct treatments such as antivirals, or through lifestyle modifications (eg, alcohol avoidance and weight loss).3Rinella ME Sanyal AJ Management of NAFLD: a stage-based approach.Nat Rev Gastroenterol Hepatol. 2016; 13: 196-205Crossref PubMed Scopus (204) Google Scholar Liver fibrosis has implications with respect to treatment initiation, follow-up, and prognosis. Greater liver scarring, without intervention, suggests a greater risk of cirrhosis and its complications long term (eg, end-stage liver failure and hepatocellular carcinoma). The challenges for clinicians include how to identify and capture populations at risk of liver disease and how to apply a test that accurately (and cheaply) identifies those individuals. Historically, liver biopsy has been regarded as the gold standard to detect liver fibrosis, but clearly, liver biopsy is not practical on a population-based level nor is it a suitable screening tool due to its invasiveness, costs, and its associated risks. Because of these reasons, there has been a proliferation of non-invasive tests of liver fibrosis over the past two decades.4Vilar-Gomez E Chalasani N Non-invasive assessment of non-alcoholic fatty liver disease: clinical prediction rules and blood-based biomarkers.J Hepatol. 2018; 68: 305-315Summary Full Text Full Text PDF PubMed Scopus (292) Google Scholar, 5Houot M Ngo Y Munteanu M Marque S Poynard T Systematic review with meta-analysis: direct comparisons of biomarkers for the diagnosis of fibrosis in chronic hepatitis C and B.Aliment Pharmacol Ther. 2016; 43: 16-29Crossref PubMed Scopus (67) Google Scholar These tests have included scores derived from routinely available blood tests and patient factors such as fibrosis-4 index (FIB-4), non-alcoholic fatty liver disease (NAFLD) fibrosis score (NFS), aspartate aminotransferase-platelet-ratio-index (APRI); panels of markers of fibrogenesis such as enhanced liver fibrosis and Fibrotest; and imaging-based tests such as transient elastography6Chang PE Goh GB Ngu JH Tan HK Tan CK Clinical applications, limitations and future role of transient elastography in the management of liver disease.World J Gastrointest Pharmacol Ther. 2016; 7: 91-106Crossref PubMed Google Scholar and magnetic resonance elastography. The imaging-based tests are accurate but expensive and not routinely available. Many, non-invasive tests seem unable to differentiate between minimal and advanced fibrosis. In short, these tests do not provide the clinician with the information they need to make informed decisions, particularly in regions where there is poor access to further tests. Screening tools are necessary in many clinical settings and it would be helpful if methods could be developed to minimise the risk of indeterminate results. It is for this reason that work in The Lancet Digital Health by Soren Sabet Sarvestany and colleagues is welcome.7Sarvestany SS Kwong JC Azhie A et al.Development and validation of an ensemble machine learning framework for detection of all-cause advanced hepatic fibrosis: a retrospective cohort study.Lancet Digit Health. 2022; 4: e188-e199Summary Full Text Full Text PDF Scopus (2) Google Scholar The researchers trained six machine learning algorithms (MLAs) using clinical information and liver biopsies from two sites in Canada (Toronto Liver Clinic, Toronto, ON, and McGill University Health Centre, Montreal, QC). The score was developed from a training set of 1703 liver biopsies and then validated in 508 biopsies across a range of fibrosis stages. The team excluded patients with decompensated cirrhosis. The authors acknowledged the controversy around the length of liver biopsy and portal tract numbers on histological interpretation,8Rockey DC Caldwell SH Goodman ZD Nelson RC Smith AD Liver biopsy.Hepatology. 2009; 49: 1017-1044Crossref PubMed Scopus (1465) Google Scholar and how the local population fibrosis stage prevalence could influence the study findings.9Poynard T Halfon P Castera L et al.Standardization of ROC curve areas for diagnostic evaluation of liver fibrosis markers based on prevalences of fibrosis stages.Clin Chem. 2007; 53: 1615-1622Crossref PubMed Scopus (208) Google Scholar The best MLA was an ensemble algorithm (ENS) of support vector machine, random forest classifier, gradient booster classifier, logistic regression, and artificial neural network. The MLA was compared against APRI, FIB-4, NAFLD-fibrosis score NFS, and transient elastography. The MLA was superior to APRI, FIB-4, and NFS across the ranges of fibrosis (AUROC curves for F0 vs F4 and F01 vs F4 the ENS ranges were 0·827 [95 CI 0·751–0·895] to 0·803 [0·733–0·863] vs 0·820 [0·743–0·891] to 0·772 [0·708–0·836] for FIB-4; 0·764 [0·678–0·850] to 0·715 [0·639–0·794] for APRI; and 0·872 [0·802–0·931] to 0·833 [0·775–0·894] vs 0·822 [0·745–0·900] to 0·758 [0·682–0·835] for NFS). Yet, the ENS was not always better than transient elastography (0·773 [0·699–0·834] vs 0·826 [0·758–0·889]) but ENS performs better than NFS. MLAs have been developed before to accurately predict liver fibrosis stages.10Hashem S Esmat G Elakel W et al.Comparison of machine learning approaches for prediction of advanced liver fibrosis in chronic hepatitis C patients.IEEE/ACM Trans Comput Biol Bioinformatics. 2018; 15: 861-868Crossref PubMed Scopus (55) Google Scholar So, what does this new MLA offer to clinicians? Indeterminate scores are frustrating to patients and clinicians alike, leading to further investigations which could be unnecessary, harmful, and cause patient anxiety. So any test which reduces this uncertainty is welcome. This model could have a role in primary care and less well-developed health care settings. Moreover, this model could minimise referrals to secondary care services and provide earlier assurances to patients. But, to be useful, the tool must be easy to use and accessible. The development of online calculators and websites has been of benefit to clinicians, and this would be necessary if this tool were to be used widely. Furthermore, the result findings must be easy to interpret (eg, correlation with fibrosis stages) if the tool is to have clinical utility. In summary, the work by Sabet Sarvestany and colleagues provides insights into how an MLA can improve the accuracy of existing models of non-invasive liver fibrosis to further refine information, reduce the number of indeterminate scores, and allow better discrimination between those patients not at risk and those at risk of significant liver scarring and cirrhosis. Whether this test becomes universally used remains to be seen, but that will be determined by how easy it is to use and interpret, its accessibility, accuracy, and its use across a spectrum of liver conditions. But this work is of interest and relevance to the field. As such, it gives a hint of the future direction, not just of liver disease but medicine more broadly. I report unrestricted educational grants from Sirtex, Bristol Myers Squibb, and Bayer; lecture fees from Dr Falk and Roche; serving on advisory boards for Bristol Myers Squibb, Sirtex, and Roche; co-writing an article with Bristol Myers Squibb; shares with AstraZeneca; and receiving a meetings pass from Roche to attend a scientific meeting. Development and validation of an ensemble machine learning framework for detection of all-cause advanced hepatic fibrosis: a retrospective cohort studyWe have shown that an ensemble MLA outperforms non-imaging-based methods in detecting advanced fibrosis across different causes of liver disease. Our MLA was superior to APRI, FIB-4, and NFS with no indeterminate classifications, while achieving performance comparable to an independent panel of experts. MLAs using routinely collected data could identify patients at high-risk of advanced hepatic fibrosis and cirrhosis among patients with chronic liver disease, allowing intervention before onset of decompensation. Full-Text PDF Open Access
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".