Sub-phenotyping of patients with severe aortic stenosis undergoing transcatheter aortic valve replacement by unsupervised agglomerative clustering of echocardiographic and hemodynamic data
Bibliographic record
Abstract
Abstract Background Severe aortic stenosis (AS) can trigger a deleterious cascade of impairments including left heart dysfunction, pulmonary hypertension (PH), and eventually right heart failure. Clinical phenotypes therefore appear heterogeneous, depending on disease progression and comorbidities. Purpose This retrospective analysis aims to categorize patients with severe AS according to clinical presentation by applying unsupervised machine learning in combination with an artificial neural network (ANN). Methods Unsupervised agglomerative clustering was applied to pre-procedural data from echocardiography and right heart catheterization from 366 consecutively enrolled patients undergoing transcatheter aortic valve replacement (TAVR) for severe AS at two tertiary centers in Germany between 2014 and 2020. Association between cluster and 2-year all-cause mortality after TAVR was assessed, and an ANN was trained to open the avenue to prospectively predict cluster assignment in future patients. Results Cluster analysis revealed four distinct phenotypes, reflecting various extents of disease severity, and hence differing in mortality. Patients from cluster 1, constituting the majority of cases and hereinafter referred to as reference, presented with regular cardiac function and without PH. Accordingly, estimated 2-year survival was 90.6% (95% CI: 85.8–95.6%). Contrarily, patients from smallest cluster 3 displayed most extensive disease characteristics, i.e. left and right heart dysfunction together with combined pre- and postcapillary PH, and their 2-year mortality was increased (2-year survival: 77.3% (95% CI: 65.2–91.6%), HR for 2-year mortality: 2.6 (95% CI: 1.1–6.2); p-value: 0.025). Clusters 2 and 4 comprised patients suffering from postcapillary PH. Whilst patients from cluster 2 showed similar survival as cluster 1 (2-year survival: 85.8% (95% CI: 76.9–95.6%)), patients from cluster 4 with right atrial enlargement and high prevalence of severe tricuspid regurgitation (TR) deceased more often (2-year survival: 74.9% (95% CI: 65.9–85.2%), HR for 2-year mortality: 2.8 (95% CI: 1.4–5.5); p-value: 0.004). After randomly dividing the study population into derivation and validation cohorts, an ANN could precisely predict cluster assignment (accuracy: 83.5%), significantly outperforming the no information rate (46.8%; p-value: 2.26e-15). Importantly, patients from high-risk clusters 3 and 4 were detected with high sensitivity (100.0% and 85.2%, respectively) and specificity (95.9% and 95.1%, respectively). Conclusion Expanding the analytical armamentarium by machine learning technology aids in capturing complex clinical presentations as observed in patients with severe AS. Assigning patients to clusters can thus facilitate a more sophisticated risk stratification in future clinical practice. Addressing irreversibility of PH and persistence of severe TR after TAVR should obtain paramount priority in order to improve long-term survival. Funding Acknowledgement Type of funding sources: Public Institution(s). Main funding source(s): Mark Lachmann receives funding from Technical University of Munich (Clinician Scientist Grant).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".