Bibliographic record
Abstract
We thank Collins and Le Manach for their insightful comments on our paper [1, 2] and also for drawing our attention to their article [1] on sample sizes for external validation of prognostic models, which unfortunately was unavailable when our data were analysed. However, their article examined the effect of sample size for proportional hazards models and we used a logistic regression model (with previously published papers [3, 4] more applicable to the assessment of appropriate sample size). Our sample size, though small, is not as exaggerated as their examples (one study with 8 cases and one with 1 case). Although we reported c-statistics with confidence intervals and the P-value for the difference between them for consistency with previously published papers, we stated upfront that this has severe limitations and that huge sample sizes would be needed to detect clinically relevant differences between two c-statistics considered to be in the ‘excellent’ range. When calculating the sample size for external validation, it is necessary to choose one or two statistics believed most important. We chose calibration; since both calibration-in-the-large and the miscalibration-coefficient were statistically significant, the power of our study is not an issue, but we acknowledge that bias of these estimates may be. Even though it does not directly apply to our logistic regression model, we did use their simulation study for a sample size of ∼37. There was no difference in the coverage rates of confidence intervals between a sample of 37 and one of 100 (or even 200). Also bias in the calibration slope is huge when the number of events is ≤10, but <2.5% when the sample size is 37 and ∼1.6% when 100. Regarding our calibration plot, the scale of the axes was chosen to avoid uninformative white space in the figure; and to avoid confusion, we added a green diagonal line to illustrate the line of equality on which should lie perfect predictions. We believe that ‘the risk of x% of how many patients died’ is clearly presented in the legend of said table and the table itself. Our calibration plots differ only from others in that both are presented on the same graph to illustrate difference. We agree that a less-smoothed calibration plot with 95% CI would have been ideal but not appropriate in this study due to small sample sizes. Operative mortality rate for all coronary artery bypass graft procedures is low (<5%); therefore, any prognostic model for operative mortality will necessarily have zero deaths in the smallest risk groups. Risk score validation of low-risk groups is equally important as for high-risk groups. Our analysis showed that calibration was strongest in the low-risk groups but in the highest risk groups was underestimated by logistic EuroSCORE and overestimated by EuroSCORE II. Surgeons can therefore be confident in either score for low-risk patients. The missing ejection fraction data are unfortunate but are currently being updated by chart review. So far, most had an ejection fraction of >50% which would not alter the EuroSCORE values nor results of our study. Finally, validation of risk models is generally accepted best assessed in the settings in which they will be used [5, 6].
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.007 | 0.069 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.003 | 0.003 |
| Scholarly communication | 0.004 | 0.006 |
| Open science | 0.003 | 0.002 |
| Research integrity | 0.046 | 0.054 |
| Insufficient payload (model declined to judge) | 0.004 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".