Cluster analysis of transcriptomic datasets to identify endotypes of Idiopathic Pulmonary Fibrosis
Bibliographic record
Abstract
ABSTRACT Background Considerable clinical heterogeneity in Idiopathic Pulmonary Fibrosis (IPF) suggests the existence of multiple disease endotypes. Identifying these endotypes could allow for a biomarker-driven personalised medicine approach in IPF. To improve our understanding of the pathogenesis of IPF by identifying clinically distinct groups of patients with IPF that could represent distinct disease endotypes. Methods We co-normalised, pooled and clustered three publicly available blood transcriptomic datasets (total 220 IPF cases). We compared clinical traits across clusters and used gene enrichment analysis to identify biological pathways and processes that were over-represented among the genes that were differentially expressed across clusters. A gene-based classifier was developed and validated using three additional independent datasets (total 194 IPF cases). Findings We identified three clusters of IPF patients with statistically significant differences in lung function (P=0·009) and mortality (P=0·009) between groups. Gene enrichment analysis implicated dysregulation of mitochondrial homeostasis, apoptosis, cell cycle and innate and adaptive immunity in the pathogenesis underlying these groups. We developed and validated a 13-gene cluster classifier that predicted mortality in IPF (high-risk clusters vs low-risk cluster: hazard ratio= 4·25, 95% confidence interval= [2·14, 8·46], P=3·7×10 −5 ). Interpretation We have identified blood gene expression signatures capable of discerning groups of IPF patients with significant differences in survival. These clusters could be representative of distinct pathophysiological states, which would support the theory of multiple endotypes of IPF. Although more work must be done to confirm the existence of these endotypes, our classifier could be a useful tool in patient stratification and outcome prediction in IPF. Funding L.V.W. holds a GSK/British Lung Foundation Chair in Respiratory Research (C17-1). R.G.J. is supported by a National Institute for Health Research (NIHR) Research Professorship (NIHR reference RP-2017-08-ST2-014). P.L.M. is supported by an Action for Pulmonary Fibrosis Mike Bray fellowship. T.M. Maher is supported by a National Institute for Health Research Clinician Scientist Fellowship (CS-2013-13-017) and a British Lung Foundation Chair in Respiratory Research (C17-3). I.N. is supported by a National Heart, Lung, and Blood Institute (NHLBI) grant (R01HL145266). D.A.S. is supported by NHLBI grants (UG3HL151865, R01HL097163, P01HL092870, X01HL134585 and UH3HL123442) and a United States Department of Defense grant (W81XWH-17-1-0597). The GSE110147 study was supported by the Roche Multi Organ Transplant Academic Enrichment Fund, Lawson Research Institute Internal Research Fund and Western Strategic Support for CIHR Success, Seed Grant. The research was partially supported by the NIHR Leicester Biomedical Research Centre; the views expressed are those of the author(s) and not necessarily those of the National Health Service (NHS), the NIHR, or the Department of Health. Putting research into context Evidence before this study We searched PubMed Central in February 2020 with the search terms “idiopathic pulmonary fibrosis”, “gene expression” and “cluster analysis” with no restrictions on publication date or language. Previous transcriptomic cluster analyses have found that differences in gene expression can be used to predict disease status, severity and outcome in IPF. A previous transcriptomic prognostic biomarker has been developed that can predict outcome in IPF using blood expression data from 52 genes. Added value of this study By utilising new methods of data co-normalisation and machine learning, we were able to combine multiple publicly available datasets and perform one of the largest transcriptomic studies in IPF to-date with a total of 416 IPF cases across the discovery and validation stages. We identified three clusters of patients, one of which appeared to contain, on average, the healthiest subjects with favourable lung function and survival over time. These clusters were defined using expression from groups of genes that were significantly enriched for many different biological pathways and processes, including metabolic changes, apoptosis, cell cycle and immune response, and so could be representative of distinct pathophysiological states. Additionally, we developed a 13-gene expression-based classifier to assign individuals with IPF to one of the clusters and validated this classifier using three additional independent cohort of IPF patients (totalling 194 IPF cases). As the clusters were associated with survival, our classifier could potentially be used to predict outcome in IPF. Implications of all the available evidence Our findings support the hypothesis that the disease consists of multiple endotypes. The clusters identified in this study could provide some valuable insight into the underlying biological processes that may be driving the considerable clinical heterogeneity in IPF. With further development, our gene expression-based classifier could be a useful tool for patient stratification and outcome prediction in IPF.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.005 | 0.004 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".