Obtaining and managing data sets for individual participant data meta-analysis: scoping review and practical guide
Bibliographic record
Abstract
BACKGROUND: Shifts in data sharing policy have increased researchers' access to individual participant data (IPD) from clinical studies. Simultaneously the number of IPD meta-analyses (IPDMAs) is increasing. However, rates of data retrieval have not improved. Our goal was to describe the challenges of retrieving IPD for an IPDMA and provide practical guidance on obtaining and managing datasets based on a review of the literature and practical examples and observations. METHODS: We systematically searched MEDLINE, Embase, and the Cochrane Library, until January 2019, to identify publications focused on strategies to obtain IPD. In addition, we searched pharmaceutical websites and contacted industry organizations for supplemental information pertaining to recent advances in industry policy and practice. Finally, we documented setbacks and solutions encountered while completing a comprehensive IPDMA and drew on previous experiences related to seeking and using IPD. RESULTS: Our scoping review identified 16 articles directly relevant for the conduct of IPDMAs. We present short descriptions of these articles alongside overviews of IPD sharing policies and procedures of pharmaceutical companies which display certification of Principles for Responsible Clinical Trial Data Sharing via Pharmaceutical Research and Manufacturers of America or European Federation of Pharmaceutical Industries and Associations websites. Advances in data sharing policy and practice affected the way in which data is requested, obtained, stored and analyzed. For our IPDMA it took 6.5 years to collect and analyze relevant IPD and navigate additional administrative barriers. Delays in obtaining data were largely due to challenges in communication with study sponsors, frequent changes in data sharing policies of study sponsors, and the requirement for a diverse skillset related to research, administrative, statistical and legal issues. CONCLUSIONS: Knowledge of current data sharing practices and platforms as well as anticipation of necessary tasks and potential obstacles may reduce time and resources required for obtaining and managing data for an IPDMA. Sufficient project funding and timeline flexibility are pre-requisites for successful collection and analysis of IPD. IPDMA researchers must acknowledge the additional and unexpected responsibility they are placing on corresponding study authors or data sharing administrators and should offer assistance in readying data for sharing.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.798 | 0.934 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.011 | 0.001 |
| Bibliometrics | 0.001 | 0.003 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.006 | 0.010 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.010 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".