Computational deconvolution of cell type-specific gene expression in COPD and IPF lungs reveals disease severity associations
Bibliographic record
Abstract
Chronic obstructive pulmonary disease (COPD) and idiopathic pulmonary fibrosis (IPF) are debilitating diseases associated with divergent histopathological changes in the lungs. At present, due to cost and technical limitations, profiling cell types is not practical in large epidemiology cohorts ( n > 1000). Here, we used computational deconvolution to identify cell types in COPD and IPF lungs whose abundances and cell type-specific gene expression are associated with disease diagnosis and severity. We analyzed lung tissue RNA-seq data from 1026 subjects (COPD, n = 465; IPF, n = 213; control, n = 348) from the Lung Tissue Research Consortium. We performed RNA-seq deconvolution, querying thirty-eight discrete cell-type varieties in the lungs. We tested whether deconvoluted cell-type abundance and cell type-specific gene expression were associated with disease severity. The abundance score of twenty cell types significantly differed between IPF and control lungs. In IPF subjects, eleven and nine cell types were significantly associated with forced vital capacity (FVC) and diffusing capacity for carbon monoxide (D L CO), respectively. Aberrant basaloid cells, a rare cells found in fibrotic lungs, were associated with worse FVC and D L CO in IPF subjects, indicating that this aberrant epithelial population increased with disease severity. Alveolar type 1 and vascular endothelial (VE) capillary A were decreased in COPD lungs compared to controls. An increase in macrophages and classical monocytes was associated with lower D L CO in IPF and COPD subjects. In both diseases, lower non-classical monocytes and VE capillary A cells were associated with increased disease severity. Alveolar type 2 cells and alveolar macrophages had the highest number of genes with cell type-specific differential expression by disease severity in COPD and IPF. In IPF, genes implicated in the pathogenesis of IPF, such as matrix metallopeptidase 7, growth differentiation factor 15, and eph receptor B2, were associated with disease severity in a cell type-specific manner. Utilization of RNA-seq deconvolution enabled us to pinpoint cell types present in the lungs that are associated with the severity of COPD and IPF. This knowledge offers valuable insight into the alterations within tissues in more advanced illness, ultimately providing a better understanding of the underlying pathological processes that drive disease progression.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".