Phenotype-based prediction of incident cardiovascular hospitalization and inpatient care costs in patients referred for cardiovascular magnetic resonance imaging: Applications of traditional statistical modelling and machine learning
Bibliographic record
Abstract
Background: Cardiovascular disease has an estimated lifetime prevalence of 48% in adults and imposes the highest economic burden on health care systems among noncommunicable diseases. These costs are largely related to chronic disease management, clinical procedures, and hospitalization, particularly for major adverse cardiovascular events (MACE). Importantly, health expenditures incurred by cardiovascular care are expected to increase substantially as the global population ages and life expectancies continue to rise. To improve health system efficiency and resource allocation in preparation for future cardiovascular care needs, it is necessary to improve baseline patient characterization and offer more accurate personalized risk predictions to optimally plan for opportunities to improve cardiovascular health while controlling costs.Aims: The aim of this thesis was to develop and validate models for the prediction of MACE and one-year cumulative inpatient care costs in a large cohort of patients referred for cardiovascular magnetic resonance imaging.Methods: Patients were recruited from the Cardiovascular Imaging Registry of Calgary, a prospective clinical outcomes registry that provides automated linkages of data abstracted from electronic health records, cardiovascular magnetic resonance imaging reports, and patient-reported health questionnaires. These data were used for predictive modelling using both traditional statistical methodologies and machine learning approaches.Results: Random survival forest and Cox proportional hazards models were developed for time-to-event prediction of hospitalization for MACE. Both models achieved time dependent AUCs of 0.83 in holdout validation. Patients with predicted risk in the upper tertile experienced 29- and 21-fold (p < 0.001) increased risk of MACE, respectively. A two-part hurdle model was developed for cost regression to predict one-year cumulative inpatient expenditures following cardiovascular magnetic resonance imaging. When binning the cost predictions into zero-, low-, and high-cost brackets, the model achieved 0.73 precision, 0.76 recall, and 0.74 F1. The best performing machine learning classification model combined predictions from random forest and artificial neural network algorithms to achieve 0.76 precision, 0.82 recall, and 0.79 F1.Conclusions: The results of this thesis demonstrate the prognostic capacity of multi-domain health data and its utility in the development of patient-specific risk models for adverse cardiovascular events and cumulative inpatient care costs. Additionally, while machine learning modelling methodologies offer advantages in handling large health care data sets, the interpretability of traditional statistical models remains valuable for delineating relationships between health-related variables and outcomes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.017 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".