Assessing the validity of cardiac surgery risk stratification systems for CABG patients in a single center.
Bibliographic record
Abstract
BACKGROUND: Most cardiac surgery risk stratification systems were primarily designed using patient-related factors to predict mortality and postoperative morbidity. Relative mortality rates are higher at cardiac surgery centers which perform surgery on elderly patients. Our aim was to assess the validity of risk stratification systems for our regional population. MATERIAL/METHODS: The study involved 1021 patients. Risk stratification was carried out using the EuroSCORE, Ontario, and QMMI scoring systems. Analysis comparing the scoring systems included sensitivity, specificity, predictive values, and receiver operating characteristics (ROC) curves. Accuracy was assessed using the systems' ability to avoid Type I and Type II errors. RESULTS: Sensitivity and specificity of the QMMI scoring system were 33.3% and 97.2%, of EuroSCORE 20.7% and 96.7%, and of Ontario 21.1% and 94.4%, respectively. The best positive predictive value was for QMMI and EuroSCORE with 75% versus Ontario's 50%. The highest negative predictive value was QMMI's 85.4% versus Ontario's 78.9% and EuroSCORE's 72.0%. The best accuracy showed QMMI scoring with 84.5% versus Ontario's 78.9% and EuroSCORE's 72.2%. CONCLUSIONS: All the investigated risk stratification systems were moderately predictive. The QMMI score showed the best predictive characteristics (sensitivity, specificity, and accuracy) for our patient population. The QMMI system had high specificity and accuracy. The EuroSCORE system showed mortality overprediction for our population, associated with high false negative test results and low accuracy. The Ontario risk stratification system often commits Type II errors, associated with a high rate of false positive test results and low accuracy.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.009 | 0.033 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".