Moving From Idealism to Realism With Data Sharing
Bibliographic record
Abstract
Ideas and OpinionsMarch 2023Moving From Idealism to Realism With Data SharingKeith A. Marsolo, PhD, Kevin P. Weinfurt, PhD, Karen L. Staman, MS, and Bradley G. Hammill, DrPHKeith A. Marsolo, PhDDepartment of Population Health Sciences, Duke University School of Medicine, Durham, North Carolina (K.A.M., K.P.W., K.L.S., B.G.H.)., Kevin P. Weinfurt, PhDDepartment of Population Health Sciences, Duke University School of Medicine, Durham, North Carolina (K.A.M., K.P.W., K.L.S., B.G.H.)., Karen L. Staman, MSDepartment of Population Health Sciences, Duke University School of Medicine, Durham, North Carolina (K.A.M., K.P.W., K.L.S., B.G.H.)., and Bradley G. Hammill, DrPHDepartment of Population Health Sciences, Duke University School of Medicine, Durham, North Carolina (K.A.M., K.P.W., K.L.S., B.G.H.).Author, Article, and Disclosure Informationhttps://doi.org/10.7326/M22-2973 SectionsAboutFull TextPDF ToolsAdd to favoritesDownload CitationsTrack CitationsPermissions ShareFacebookTwitterLinkedInRedditEmail Significant efforts have been made in the past decade to promote open science and data sharing in clinical research. The moral and scientific arguments are clear: If data are shared, it could promote transparency and understanding of the results, honor the participation of individuals, and enable new discoveries (1).The White House Office of Science and Technology Policy recently updated guidance requiring that results of federally funded research be made immediately available, and federal agencies have drafted a series of policies that outline expectations of their awardees. For example, the National Institutes of Health (NIH) has released a new Policy ...References1. The benefits of data sharing. In: Institute of Medicine. Sharing Clinical Research Data: Workshop Summary. National Academies Pr; 2013. Accessed at www.ncbi.nlm.nih.gov/books/NBK137823 on 9 May 2022. Google Scholar2. National Institutes of Health. Final NIH Policy for Data Management and Sharing. 2022. Accessed at https://grants.nih.gov/grants/guide/notice-files/NOT-OD-21-013.html on 11 November 2022. Google Scholar3. Patient-Centered Outcomes Research Institute. Policy for Data Management and Data Sharing. 2018. Accessed at www.pcori.org/about-us/governance/policy-data-management-and-data-sharing on 13 May 2022. Google Scholar4. National Institutes of Health. Supplemental information to the NIH Policy for Data Management and Sharing: protecting privacy when sharing human research participant data. 2022. Accessed at https://grants.nih.gov/grants/guide/notice-files/NOT-OD-22-213.html on 12 December 2022. Google Scholar5. Wilkinson MD, Dumontier M, Aalbersberg IJ, et al. The FAIR Guiding Principles for scientific data management and stewardship. Sci Data. 2016;3:160018. [PMID: 26978244] doi:10.1038/sdata.2016.18 CrossrefMedlineGoogle Scholar6. European Medicines Agency. Clinical data publication. 2018. Accessed at www.ema.europa.eu/en/human-regulatory/marketing-authorisation/clinical-data-publication on 19 November 2022. Google Scholar7. Gøtzsche PC, Jørgensen AW. Opening up data at the European Medicines Agency [Letter]. BMJ. 2011;342:d2686. [PMID: 21558364] doi:10.1136/bmj.d2686 CrossrefMedlineGoogle Scholar8. Herder M, Doshi P, Lemmens T. Precedent pushing practice: Canadian court orders release of unpublished clinical trial data. BMJ Opinion. 19 July 2018. Accessed at https://blogs.bmj.com/bmj/2018/07/19/precedent-pushing-practice-canadian-court-orders-release-of-unpublished-clinical-trial-data on 19 November 2022. Google Scholar9. National Institutes of Health. NIH Genomic Data Sharing Policy. 2014. Accessed at https://grants.nih.gov/grants/guide/notice-files/not-od-14-124.html on 12 December 2022. Google Scholar Author, Article, and Disclosure InformationAuthors: Keith A. Marsolo, PhD; Kevin P. Weinfurt, PhD; Karen L. Staman, MS; Bradley G. Hammill, DrPHAffiliations: Department of Population Health Sciences, Duke University School of Medicine, Durham, North Carolina (K.A.M., K.P.W., K.L.S., B.G.H.).Disclaimer: The content is solely the responsibility of the authors and does not necessarily represent the official views of the NIH or its HEAL Initiative.Financial Support: This work is supported within the NIH Health Care Systems Research Collaboratory by the NIH Common Fund through cooperative agreement U24AT009676 from the Office of Strategic Coordination within the Office of the NIH Director. This work is also supported by the NIH through the NIH HEAL Initiative under award U24AT010961.Disclosures: Disclosures can be viewed at www.acponline.org/authors/icmje/ConflictOfInterestForms.do?msNum=M22-2973.Corresponding Author: Keith A. Marsolo, PhD, Population Health Sciences, Duke University, 300 West Morgan Street, Suite 636, Durham, NC 27701; e-mail, keith.marsolo@duke.edu.Author Contributions: Conception and design: B.G. Hammill, K.A. Marsolo, K.P. Weinfurt.Analysis and interpretation of the data: K.L. Staman.Drafting of the article: B.G. Hammill, K.A. Marsolo, K.L. Staman, K.P. Weinfurt.Critical revision for important intellectual content: B.G. Hammill, K.A. Marsolo, K.P. Weinfurt.Final approval of the article: B.G. Hammill, K.A. Marsolo, K.L. Staman, K.P. Weinfurt.Obtaining of funding: K.P. Weinfurt.This article was published at Annals.org on 31 January 2023. PreviousarticleNextarticle Advertisement FiguresReferencesRelatedDetails Metrics March 2023Volume 176, Issue 3Page: 402-403KeywordsAlgorithmsData managementDisclosureHealth careHealth Insurance Portability and Accountability ActReproducibilityResearch fundingScience policyStatistical dataStatistical methods ePublished: 31 January 2023 Issue Published: March 2023 Copyright & PermissionsCopyright © 2023 by American College of Physicians. All Rights Reserved.PDF downloadLoading ...
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.211 | 0.243 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.009 | 0.026 |
| Scholarly communication | 0.023 | 0.019 |
| Open science | 0.004 | 0.018 |
| Research integrity | 0.012 | 0.030 |
| Insufficient payload (model declined to judge) | 0.018 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".