MétaCan
Menu
Back to cohort
Record W4388500358 · doi:10.1093/eurheartj/ehad718

Endpoint adjudication in cardiovascular clinical trials

2023· article· en· W4388500358 on OpenAlexaff
Muhammad Shahzeb Khan, Muhammad Usman, Harriette G.C. Van Spall, Stephen J. Greene, Omar Baqal, G. Michael Felker, Deepak L. Bhatt, James L. Januzzi, Javed Butler

Bibliographic record

VenueEuropean Heart Journal · 2023
Typearticle
Languageen
FieldMathematics
TopicStatistical Methods in Clinical Trials
Canadian institutionsMcMaster UniversitySt. Joseph’s Healthcare HamiltonImpact
Fundersnot available
KeywordsMedicineAdjudicationClinical endpointSurrogate endpointEndpoint DeterminationUnstable anginaClinical trialMyocardial infarctionBlindingGeneralizability theoryRandomized controlled trialIntensive care medicineCardiologyInternal medicineStatistics

Abstract

fetched live from OpenAlex

Endpoint adjudication (EA) is a common feature of contemporary randomized controlled trials (RCTs) in cardiovascular medicine. Endpoint adjudication refers to a process wherein a group of expert reviewers, known as the clinical endpoint committee (CEC), verify potential endpoints identified by site investigators. Events that are determined by the CEC to meet pre-specified trial definitions are then utilized for analysis. The rationale behind the use of EA is that it may lessen the potential misclassification of clinical events, thereby reducing statistical noise and bias. However, it has been questioned whether this is universally true, especially given that EA significantly increases the time, effort, and resources required to conduct a trial. Herein, we compare the summary estimates obtained using adjudicated vs. non-adjudicated site designated endpoints in major cardiovascular RCTs in which both were reported. Based on these data, we lay out a framework to determine which trials may warrant EA and where it may be redundant. The value of EA is likely greater when cardiovascular trials have nuanced primary endpoints, endpoint definitions that align poorly with practice, sub-optimal data completeness, greater operator variability, and lack of blinding. EA may not be needed if the primary endpoint is all-cause death or all-cause hospitalization. In contrast, EA is likely merited for more nuanced endpoints such as myocardial infarction, bleeding, worsening heart failure as an outpatient, unstable angina, or transient ischaemic attack. A risk-based approach to adjudication can potentially allow compromise between costs and accuracy. This would involve adjudication of a small proportion of events, with further adjudication done if inconsistencies are detected.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.220
metaresearch head score (Gemma)0.464
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Insufficient payload (model declined to judge)
Consensus categoriesMetaresearch
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Theoretical or conceptual · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.563
Threshold uncertainty score0.999

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.2200.464
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0020.001
Bibliometrics0.0000.001
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0000.002

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.898
GPT teacher head0.672
Teacher spread0.226 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designTheoretical or conceptual
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations11
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueEuropean Heart JournalSame topicStatistical Methods in Clinical TrialsFrench-language works237,207