MétaCan
Menu
Back to cohort
Record W2241662783 · doi:10.1093/ejcts/ezv190

Knowingly repeating an incorrect and inefficient analysis is flawed logic

2015· letter· en· W2241662783 on OpenAlexaff
Gary S. Collins, Yannick Le Manach

Bibliographic record

VenueEuropean Journal of Cardio-Thoracic Surgery · 2015
Typeletter
Languageen
FieldComputer Science
TopicMachine Learning in Healthcare
Canadian institutionsMcMaster UniversityPopulation Health Research Institute
Fundersnot available
KeywordsComputer scienceBusiness

Abstract

fetched live from OpenAlex

In their study, Garcia-Valentin et al. [ 1 ] evaluated the performance of EuroSCORE and EuroSCORE II in a cohort of 4034 patients undergoing cardiac surgery in Spain. While we applaud the authors for conducting this independent external validation, which is uncommon, we wish to highlight some methodological and reporting concerns so that future investigators do not replicate them. The two key components characterizing the performance of a prediction model are discrimination and calibration. Discrimination is the ability of the prediction model to differentiate between those who do and do not experience the outcome event (quantified by c -index which is equivalent to the area under the receiver operating characteristic curve). Calibration is the agreement between outcome predictions from the model and the observed outcomes. In the study by Garcia-Valentin et al. , calibration was evaluated using the Hosmer–Lemeshow test, which the authors correctly highlight as being problematic. The Hosmer–Lemeshow test has limited power to evaluate calibration, affected by sample size and grouping and gives no indication on the direction and magnitude of (mis)calibration [ 2 ]. Unfortunately, the authors disappointingly proceeded to use this test to judge calibration on the premise that this was used in the original model development study [ 3 ]. Repeating an incorrect and uninformative analysis on the grounds that it was done in the original study is flawed logic and if methodological concerns are raised, then alternative and correct analyses should be carried out. Knowingly repeating a flawed analysis even if concerns are raised, incorrectly supports and justifies its use by other investigators and the cycle is never broken. The study by Garcia-Valentin et al. raised concerns of the approach but only in the ‘Discussion’ section, which is often not read as closely as the ‘Methods’ or ‘Results’ section of a manuscript. In accordance with recent recommendations on the reporting of prediction model studies [ 2 , 4 ], calibration should be assessed graphically by plotting predicted outcome probabilities ( x -axis) against observed outcomes ( y -axis) using a high-resolution smoothed (loess) line. The direction and magnitude of any miscalibration can then be examined across the entire probability range. The calibration plot can also be supplemented with a numerical quantification of calibration by examining the calibration slope and intercept and the 0.9 quantile of the absolute prediction error [ 5 , 6 ]. We recommend investigators, peer reviewers and editors to read the recent guidance from the Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis (TRIPOD) Initiative ( www.tripod-statement.org ). The TRIPOD reporting guideline for clinical prediction models, discusses key issues in the development and validation of a prediction model [ 4 ]. The TRIPOD guideline is similar to other well-known reporting guidelines (e.g. CONSORT, STROBE and PRISMA) designed to help authors, peer reviewers and journal editors in ensuring that the essential items describing the development or validation of a clinical prediction model are clearly reported. Accompanying the reporting guideline is an extensive Explanation and Elaboration article describing the rationale for the checklist item but also highlighting many methodological considerations when developing or validating a clinical prediction model [ 2 ].

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.049
metaresearch head score (Gemma)0.314
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.951
Threshold uncertainty score0.257

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0490.314
Meta-epidemiology (narrow)0.0010.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0010.001
Science and technology studies0.0010.006
Scholarly communication0.0030.004
Open science0.0020.001
Research integrity0.0110.015
Insufficient payload (model declined to judge)0.0030.006

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.058
GPT teacher head0.317
Teacher spread0.259 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
DomainMethods
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations1
Published2015
Admission routes1
Has abstractyes

Explore more

Same venueEuropean Journal of Cardio-Thoracic SurgerySame topicMachine Learning in HealthcareFrench-language works237,207