Medical journal requirements for clinical trial data sharing: Ripe for improvement
Bibliographic record
Abstract
In some science, technology, engineering, and mathematics (STEM) fields, data sharing is the norm (e.g., physics or space science).However, this is currently not the case in biomedicine, except for certain exceptions in areas such as genomics.For therapeutic research, data sharing is expected to maximize the value of research for clinical practice by means of greater transparency and opportunities for external researchers to reanalyze, synthesize, replicate, and build upon previous evidence.Examples include reanalyses, secondary analyses, individual patient data (IPD) meta-analyses, and methodological evaluations.Maximizing the efficient use of clinical research data is important in the development of new therapeutic options, including treatments for the Coronavirus Disease 2019 (COVID-19). Summary points• Efficient sharing and reuse of data from clinical trials are critical in advancing medical knowledge and developing improved treatments.• We believe that the International Committee of Medical Journal Editors (ICMJE) clinical trial data sharing policy is currently inadequate.• Although data sharing plans help increase transparency, they do not ensure that data are shared, and they are often inadequately implemented.• We believe that the ICMJE should adapt a stronger policy on data sharing that is enforced rigorously in all ICMJE members and affiliated journals.• The policy should include a strong evaluation component to ensure that all clinical trial data are shared, their value maximized, and data producers incentivized.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.790 | 0.913 |
| Meta-epidemiology (narrow) | 0.002 | 0.008 |
| Meta-epidemiology (broad) | 0.010 | 0.012 |
| Bibliometrics | 0.016 | 0.031 |
| Science and technology studies | 0.007 | 0.014 |
| Scholarly communication | 0.038 | 0.047 |
| Open science | 0.017 | 0.023 |
| Research integrity | 0.028 | 0.028 |
| Insufficient payload (model declined to judge) | 0.071 | 0.050 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".