MétaCan
Menu
Back to cohort
Record W2282129424 · doi:10.1136/jech-2015-206698

Challenges in reproducing results from publicly available data: an example of sexual orientation and cardiovascular disease risk

2016· article· en· W2282129424 on OpenAlexaff
Nichole Austin, Sam Harper, Jay S. Kaufman, Ghassan B. Hamra

Bibliographic record

VenueJournal of Epidemiology & Community Health · 2016
Typearticle
Languageen
FieldMedicine
TopicSex and Gender in Healthcare
Canadian institutionsMcGill University
Fundersnot available
KeywordsNational Health and Nutrition Examination SurveyMedicineCovariateSexual orientationReplication (statistics)ReplicateNull hypothesisType I and type II errorsStatisticsFramingham Heart StudyEconometricsDiseaseFramingham Risk ScorePopulationEnvironmental healthPathologyPsychologySocial psychology

Abstract

fetched live from OpenAlex

BACKGROUND: Replication is a vital part of the research process and has recently received considerable attention. Analyses using publicly available data should, if adequately described, be reproducible without assistance from the original investigators. Using data from the US National Health and Nutrition Examination Survey (NHANES), a recent study reported a statistically significant difference in cardiovascular disease risk comparing subgroups of sexual minority men. We attempted to reproduce these findings and assessed whether the results were robust to alternative analytic strategies and assumptions. METHODS: We used the exclusion criteria and coding strategy described in the original paper to construct our analytical data set. Sampling weights were constructed in accordance with NHANES analytical guidelines. We estimated crude and covariate-adjusted associations between sexual orientation and vascular age using the regression models specified in the original report. We also conducted a series of sensitivity analyses to improve on the original findings. RESULTS: Our replication attempt was partially successful: we replicated the general trends reported in the original analysis, but not identical effect estimates. Importantly, we identified a potential misapplication of the Framingham Risk Score; correcting for this increased the probability that the reported null hypothesis test was a type I error. CONCLUSIONS: This paper supports the recent calls for greater transparency and improved reporting in research. Even with a publicly available and well-documented data source, we were unable to exactly replicate another study's original findings. Our sensitivity analyses revealed key issues in the original analysis and demonstrate the scientific importance of research replication.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Direct model labels (unvalidated)

Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.

Model armCategoriesStudy designConfidence
gemmaMetaresearch
Domain: Reproducibility · Genre: Empirical
About the Canadian research system: no · About a Canadian topic: no
Observationallow
gptMetaresearch
Domain: Reproducibility · Genre: Empirical
About the Canadian research system: no · About a Canadian topic: no
Other designhigh
models splitAgreement compares identical category sets and study designs across arms.

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.788
metaresearch head score (Gemma)0.914
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Reproducibility · Consensus signal: none
Study designCandidate signal: Observational · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.212
Threshold uncertainty score0.261

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.7880.914
Meta-epidemiology (narrow)0.0030.003
Meta-epidemiology (broad)0.0040.009
Bibliometrics0.0090.014
Science and technology studies0.0060.019
Scholarly communication0.0130.012
Open science0.0100.012
Research integrity0.0100.016
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.617
GPT teacher head0.453
Teacher spread0.164 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Labeled directly by 2 models reading the full record.

Metaresearch

The models disagree on parts of this classification; every voice is preserved in the section at the end of the page.

Study designObservational · Other design
DomainReproducibility
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations3
Published2016
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Epidemiology & Community HealthSame topicSex and Gender in HealthcareCategoryMetaresearchFrench-language works237,207