MétaCan
Menu
Back to cohort
Record W4318617680 · doi:10.29220/csam.2023.30.1.001

A case study of competing risk analysis in the presence of missing data

2023· article· en· W4318617680 on OpenAlexafffund
Limei Zhou, Peter C. Austin, Husam Abdel-Qadira

Bibliographic record

VenueCommunications for Statistical Applications and Methods · 2023
Typearticle
Languageen
FieldMathematics
TopicStatistical Methods and Inference
Canadian institutionsSunnybrook HospitalUniversity of TorontoInstitute for Clinical Evaluative Sciences
FundersOntario Ministry of Health and Long-Term CareCancer Care OntarioHeart and Stroke Foundation of Canada
KeywordsMissing dataEconometricsStatisticsMathematics

Abstract

fetched live from OpenAlex

Observational data with missing or incomplete data are common in biomedical research.Multiple imputation is an effective approach to handle missing data with the ability to decrease bias while increasing statistical power and efficiency.In recent years propensity score (PS) matching has been increasingly used in observational studies to estimate treatment effect as it can reduce confounding due to measured baseline covariates.In this paper, we describe in detail approaches to competing risk analysis in the setting of incomplete observational data when using PS matching.First, we used multiple imputation to impute several missing variables simultaneously, then conducted propensity-score matching to match statin-exposed patients with those unexposed.Afterwards, we assessed the effect of statin exposure on the risk of heart failure-related hospitalizations or emergency visits by estimating both relative and absolute effects.Collectively, we provided a general methodological framework to assess treatment effect in incomplete observational data.In addition, we presented a practical approach to produce overall cumulative incidence function (CIF) based on estimates from multiple imputed and PS-matched samples.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.023
metaresearch head score (Gemma)0.061
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.023
Threshold uncertainty score0.119

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0230.061
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0020.003
Bibliometrics0.0020.003
Science and technology studies0.0020.003
Scholarly communication0.0040.003
Open science0.0020.003
Research integrity0.0060.004
Insufficient payload (model declined to judge)0.0080.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.423
GPT teacher head0.581
Teacher spread0.158 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2023
Admission routes2
Has abstractyes

Explore more

Same venueCommunications for Statistical Applications and MethodsSame topicStatistical Methods and InferenceFrench-language works237,207