Predicting outcomes in cardiac surgery: risk stratification matters?
Bibliographic record
Abstract
PURPOSE OF REVIEW: To illustrate the limitations of predictive risk models in cardiac surgery, highlight the difficulty in interpreting risk-adjusted outcome analysis and discuss the challenges of making clinical decisions based on risk predictions, particularly in high-risk patients. RECENT FINDINGS: Predictive risk models developed after logistic regression or other complex statistical analysis are commonly perceived as rigorous means to determine risk-adjusted mortality in cardiac surgery. However, the discrimination provided by those predictive models is barely better than clinical judgment. Moreover, validation studies of those models show that their calibration is inconsistent, limiting their application for comparisons between different patient cohorts. Recent data also show that, without a reasonable overlap of case-mix distributions, apparently calibrated models used for risk-adjusted outcome analysis may lead to inaccurate side-by-side comparisons of provider performance. Finally, most predictive models overestimate risk, particularly in the high-risk patients. SUMMARY: Failure to account for many biological and procedural variables and for the constantly evolving practice of surgery and perioperative medicine likely contributes to the modest predictive performance of risk models in cardiac surgery. Consequently, those models should have limited input in the analysis of provider performance and in the decision to accept or deny surgery to the high-risk patients.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.005 | 0.005 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".