Practice of data sharing plans in clinical trial registrations and concordance between registered and published data sharing plans: a cross-sectional study
Bibliographic record
Abstract
BACKGROUND: The International Committee of Medical Journal Editors (ICMJE) recommends that trial authors must specify data sharing plans when trials are registered and published, yet this uptake remains unclear. We aimed to assess the practice of data sharing plans in trial registration platforms and the concordance between registered and published data sharing plans. METHODS: We included clinical trials published between 2021 and 2023 in six high-profile journals (The Lancet, The New England Journal of Medicine, JAMA, BMJ, JAMA Internal Medicine, and Annals of Internal Medicine) that enrolled participants no earlier than 2019 and registered on clinical trial platforms. One study outcome was data sharing plans in the trial registration platform, where trials clearly responding a "yes" to "Plan to share" were considered as planning to share data (including study protocols, statistical analysis plans, analytic codes, and individual participant data). The concordance between registered and published plans to share data was also assessed, which included plans to either share data (Yes/Yes) and not to share data (No/No) in both registration and publications. Univariate analyses were used to assess associations between trial characteristics and registered plans to share data and between trial characteristics and concordance. RESULTS: Of the 383 included registration IDs, only 44.6% (171/383) planned to share data in registration. Trials with drug versus non-drug interventions had increased odds of registering plans to share data (OR = 2.71, 95% CI: 1.63, 4.63). There were seven trial publications, each pooling two trials and having two registration IDs. We selected the registration IDs with a later start date, resulting in 376 trial publications for concordance assessment. Over half (216/376, 57.4%) had discordance between registration and publications. COVID-19-related trials were associated with decreased odds of data sharing concordance (OR = 0.59, 95% CI: 0.37, 0.91). Additionally, significant discordance was consistently found in statistical analysis plans or study protocols, analytic codes, and individual participant data. CONCLUSIONS: Most registered trials do not specify plans to share data. More than half of published trials have data sharing discordance between registration and publication. Efforts are required to improve the reporting and reliability of plans to share clinical trial data. TRIAL REGISTRATION: This study was registered on the Open Science Framework ( https://osf.io/k6etb ).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | MetaresearchOpen science Domain: Reproducibility · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Observational | high |
| gpt | MetaresearchOpen science Domain: Reporting · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Observational | high |
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.048 | 0.331 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.000 | 0.002 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".