Imaging-based clusters in former smokers of the COPD cohort associate with clinical characteristics: the SubPopulations and intermediate outcome measures in COPD study (SPIROMICS)
Bibliographic record
Abstract
BACKGROUND: Quantitative computed tomographic (QCT) imaging-based metrics enable to quantify smoking induced disease alterations and to identify imaging-based clusters for current smokers. We aimed to derive clinically meaningful sub-groups of former smokers using dimensional reduction and clustering methods to develop a new way of COPD phenotyping. METHODS: An imaging-based cluster analysis was performed for 406 former smokers with a comprehensive set of imaging metrics including 75 imaging-based metrics. They consisted of structural and functional variables at 10 segmental and 5 lobar locations. The structural variables included lung shape, branching angle, airway-circularity, airway-wall-thickness, airway diameter; the functional variables included regional ventilation, emphysema percentage, functional small airway disease percentage, Jacobian (volume change), anisotropic deformation index (directional preference in volume change), and tissue fractions at inspiration and expiration. RESULTS: We derived four distinct imaging-based clusters as possible phenotypes with the sizes of 100, 80, 141, and 85, respectively. Cluster 1 subjects were asymptomatic and showed relatively normal airway structure and lung function except airway wall thickening and moderate emphysema. Cluster 2 subjects populated with obese females showed an increase of tissue fraction at inspiration, minimal emphysema, and the lowest progression rate of emphysema. Cluster 3 subjects populated with older males showed small airway narrowing and a decreased tissue fraction at expiration, both indicating air-trapping. Cluster 4 subjects populated with lean males were likely to be severe COPD subjects showing the highest progression rate of emphysema. CONCLUSIONS: QCT imaging-based metrics for former smokers allow for the derivation of statistically stable clusters associated with unique clinical characteristics. This approach helps better categorization of COPD sub-populations; suggesting possible quantitative structural and functional phenotypes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".