MétaCan
Menu
← Back to cohort
Record W4386910423 · doi:10.1681/asn.0000000000000223

The Collaborative Nephrology Community: Perspectives and Experience on Data Sharing

2023· article· en· W4386910423 on OpenAlexfundno aff
Lesley A. Inker, Juhi Chaudhari, Tom Greene, Anthony Gucciardo, Hiddo J.L. Heerspink

Bibliographic record

VenueJournal of the American Society of Nephrology · 2023
Typearticle
Languageen
FieldMedicine
TopicChronic Kidney Disease and Diabetes
Canadian institutionsnot available
FundersJanssen PharmaceuticalsNational Center for Advancing Translational SciencesNovo NordiskBayer FundAstraZenecaCSL BehringAmerican Society of NephrologyNational Heart, Lung, and Blood InstituteNovartis Pharmaceuticals CorporationGlaxoSmithKlineAstraZeneca United StatesNational Institute of Diabetes and Digestive and Kidney DiseasesChinook Therapeutics
KeywordsData sharingClinical trialMedicineHealth Insurance Portability and Accountability ActEuropean unionAccountabilityPublic healthPublic relationsBusinessFamily medicinePolitical scienceHealth careInternal medicineAlternative medicineNursingLawPathology

Abstract

fetched live from OpenAlex

Data from clinical trials or observational studies shared with the purpose of individual patient meta-analyses have transformed kidney disease clinical practice, research, and public health. We, in the CKD Epidemiology Collaboration Clinical Trials (CKD-EPI CTs), have been involved in data sharing through secondary analyses of clinical trials.1 Using these shared resources, we have provided evidence to support validation and use of GFR decline and changes in albuminuria as surrogate end points for CKD progression.2,3 In this perspective, we describe the current state of data sharing and our experience and challenges and propose ways for further improvement. Over the past 25 years, formalized mechanisms for data sharing have grown exponentially. The National Institutes of Health first developed data and sample repositories, followed by the emergence and subsequent growth of data sharing platforms (DSPs) for industry-sponsored trials.4 In parallel, laws were implemented to protect patient privacy, such as the US Health Insurance Portability and Accountability Act (1996), General Data Protection Regulation for European Union Member States (2018), and Personal Information Protection Law in the People's Republic of China (2021). In follow-up to the International Committee of Medical Journal Editors' statement of our “ethical obligation to share data because trial participants have put themselves at risk,” their 2017 requirement for sharing of individual participant data in order for consideration for publication solidified data sharing as a vital component of all dissemination activities of all trial results.5 Funding for CKD-EPI CTs began as a collaborative agreement with the National Institute of Diabetes and Digestive and Kidney Diseases in 2003 and continued work through periodic support for the purposes of scientific workshops sponsored by the National Kidney Foundation in collaboration with the Food and Drug Administration and European Medicines Agencies. In 2018, through the support of the National Kidney Foundation as our administrative core, we created a standing research group that allowed enhancement of our data coordination systems, data sharing practices, and analytical methods, resulting in further progress in advancing the evidence to support surrogate end points. We now offer our academic investigators who shared data and sponsors with a return in sharing through distribution of preliminary and confidential results, opportunities to be included in Writing Groups for manuscripts, opportunities to contribute in discussion of implications and future of analyses, and opportunities for sponsors to ask data-related questions or for academic collaborators to request permission to conduct ancillary studies using the data. In addition, we post our analytical code on public sites.6 Before 2018, we had 47 randomized controlled trials; all but two were received through agreements with academic collaborators. For most of that era, there were no formal mechanism to request individual patient data; our knowledge of the field, connections to collaborators, and willingness to work beyond available resources were key to our ability to obtain access to these data. Times are different, both regarding the formal avenues to request data and the number of larger trials evaluating CKD progression. Since 2018, we added 299 randomized controlled trials; 16 were received directly from sponsors and seven from DSP that did not require any approval other than application to the DSP. Despite advances in, and enthusiasm for, data sharing, acquiring and analyzing data is often a slower, more complex process. Indeed, for both before and since 2018, we were not able to receive data from approximately one-third of identified studies. Our knowledge, connections, and intense efforts remain critical to our data sharing process. Table 1 summarizes steps and processes in acquisition and use of data sharing. We highlight selected challenges below. Data source: Once a study is identified from our systematic review or knowledge of the field, it can be a challenge to determine how the data can be accessed. Published data sharing statements do not necessarily indicate how to request permission for data. Even when trials are shared on DSP platforms, lack of interoperability or integration across the several platforms makes it a challenge to locate studies. Approvals: We often require approval from sponsors and steering committees to obtain data because both often have rights to data through steering or publication committee charters. The latter do not necessarily approve the request even with sponsor approval, and we find that our relationships with steering committee members greatly facilitate this process. Terms of agreements: CKD-EPI CTs developed a common agreement across the three CKD-EPI CT institutions to facilitate negotiation. Despite this, several iterations are often required before the agreement can be finalized. At the extreme, in one case, it took 2 years from initial verbal agreement to signed agreement and another 12 months to receive the data. Agreements that take even a year to finalize would preclude use of such data in a time-sensitive project, such as analyses to support scientific workshops or R01 grants. In our experience, the issues raised are not salient to the work, including extreme protection of patenting or proprietary rights. Anonymization: The data privacy laws vary in the nature and magnitude of information suppression required, thus creating challenges in sharing trial data from multinational trials. For example, in the United States, deidentified data are sufficient for sharing but not so for General Data Protection Regulation or Personal Information Protection Law. Data anonymization strategies implemented to share data adherent to these laws may prevent conduct of certain analyses. For example, suppression of age or sex impedes calculation of eGFR, a required variable in our analyses. To date, we have successfully negotiated with most sponsors regarding anonymization strategies, but this is not possible for data shared on DSP. More importantly, increasingly strong protection laws may preclude use of subsets or whole studies or create datasets that cannot provide sufficiently detailed information for use in analyses. Quality control: Our highly engaged group of statisticians, domain experts, and information technology experts has developed a process for data cleaning, with multiple checkpoints for quality control. We ensure consistency across studies and of each study with published results; variations in both could bias our results. For example, in one large study, we observed a substantial difference in the treatment effect between our analyses vs the published results. After several months of discussions with the investigators, we resolved the discrepancy by adding in an adjustment factor in our statistical models to account for the hierarchical study design. It is hence crucial that study sponsors document the analysis plan, with variables specified where possible, and make this documentation available along with the data sharing package. Often study coordinating centers are disbanded, making it difficult to resolve questions. In one National Institutes of Health study, we are unable to resolve inconsistent findings between our computed kidney failure end point and the study provided end point, also inconsistent with the observed GFR decline, precluding use of this study. Meta-analyses: We continually add studies to our trial library and refine our codes to perform analyses within studies and meta-analyses across the studies or subgroups of the studies. Both sponsors and DSP agreements often impose time restrictions, such as the requirement to publish the results of the analysis within 1 year of analysis. These time lines may be insufficient for comprehensive analyses and final acceptance in journals and limit inclusion of earlier studies in updated analyses performed as new studies are obtained and new methods are developed, but access to prior studies is lost. Table 1 - Key steps and processes in data sharing with challenge and solutions Step Processes Challenge Solution Data acquisition Study identification Determination whether the study meets eligibility criteria before data acquisition Reporting of components of composite end points Data source Challenge in locating the study on DSP Increase transparency of DSP by integration and interoperability across DSP with increased participation from academics investigators and biotech companies Data agreements Approvals Sponsors concern of adherence to privacy laws and disclosure of proprietary information to competitors or before regulatory approval, and indicate that they may agree but at some point in the future Making data available after regulatory authorization of a drug Steering committee does not agree to data sharing because of lack of protection of academic interests or concern that erroneous results will be in the public domain A priori steering committee charter that includes release of data to DSP within a fixed period after primary result publicationFormalize inclusion of a representative study investigator as the co-author on resulting publications; maintain communication even if the study was acquired through DSP Terms of agreements Several iterations required over 3–12 mo before final agreement Standardized agreements with prespecified choices for the contentious items delineated, limiting the need for negotiation Data analyses Analyzing data through DSP Computation speed and storage capacities of the host; cost; academic institution firewall may prevent access Consolidated or integration and interoperability of DSP allowing for increased computing resources Anonymization Anonymization prevents computation of variables or analyses of subgroups or inability to verify published results Advance anonymization standards to accommodate more dynamic rule structures Inability to replicate published results because protection laws preclude use of a subset of studies Provide computed treatment effects of shared data Quality control Inadequate information on variables Clinical Data Interchange Standards Consortium that provide common frameworks and terminology for study dataComprehensive data package, including protocol, data dictionary, computed treatment effects if the data have been reduced or anonymized, and code to re-create the main analysis Meta-analysis across studies Inability to include earlier studies in updated analyses because of expired agreements Agreements to terminate at the end of research completion Dissemination of results Peer-reviewed publications Inability to adhering to data sharing requirements by journals because of agreements Refining ICMJE requirements for analyses of third-party data Inability to provide requested analyses because of expired agreements or accessing data again is too costly or time consuming Direct communications between authors and editorsAgreements to terminate at the end of research completionNo or reduced cost associated with reply to reviewers DSP, data sharing platform; ICMJE, International Committee of Medical Journal Editor. All solutions must be attentive to the concerns of all stakeholders, such as privacy for patients, use of proprietary information by competitors or unfavorable regulatory decisions for sponsors, loss of control by sponsors and steering committees for results in the public domain, or insufficient time to publish results of investigators. Evidence suggests that even these parties appreciate the benefits of data sharing. In a survey of clinical trial participants, few had strong concerns about the risks of data sharing,7 and sponsors have been strong supporters of our work in CKD-EPI CTs, recognizing that our analyses can help inform the design of their next trials. Selected proposed solutions: First, data standards, such as the Clinical Data Interchange Standards Consortium, that provide common frameworks and terminology for study data should be required. Second, when data are shared, the sponsors or investigators should provide a comprehensive data package, including protocol, data dictionary, computed treatment effects if the data have been reduced or anonymized, and software code to re-create the main analysis. Third, we support increased transparency of DSP with enhanced computing resources,4 and increased participation from academic investigators and biotech companies. Fourth, we recommend standardized agreements with prespecified choices for the contentious items delineated, limiting the need for negotiation. Fifth, we encourage international collaborations to harmonize data sharing standards and accommodation of more dynamic rule structures for anonymization.8 None of these concerns are trivial. The first two suggestions can be implemented immediately without discussion with other entities. The third and fourth will require cooperation by many stakeholders, including study sponsors, steering committees, representatives from DSPs, and consumers of secondary data. The fifth will require cooperation by such stakeholders together with government representatives led by highly dedicated champions. Costs are substantial and distributed across several institutions, making it nearly impossible to quantify the total amount spent on data sharing. There are costs for sponsors, funding agencies, or investigators for preparation of the data to be shared along with supporting documents. There are also costs for maintenance of the sponsor-specific platforms or DSPs. The latter are maintained by large organizations, which require expensive human and computing resources. Use of shared data by research groups, such as ours, also requires extensive support for our infrastructure to identify, acquire, clean, and ultimately maintain the data in a secured setting; perform complex analyses across multiple shared studies; and preserve the relationships that are at the heart of our collaboration. Despite the overall high costs, there remains no entity responsible to address questions related to data once shared. Finally, we suggest that journals appreciate that data sharing requirements cannot be applied to analyses of third-party data. Analyses suggested by reviewers often are not possible given the data structure, availability, or costs. Direct communication between editors and authors may help to prioritize analyses and enhance the resulting publications. Our close and collaborative nephrology community allowed us to use individual participant data to address critical questions in CKD epidemiology well before the introduction of formalized mechanisms in data sharing. Our success over the past two decades has been because of our dedication to the questions, perseverance, and utmost respect for the participants and investigators who generated the data; our ability to bring together advocacy groups, regulators, sponsors, and academic investigators; and, most importantly, the strong support from our collaborators. These factors might not be sufficient to overcome all of the challenges we face as we enter our third decade. Other research consortia without ample experience, knowledge of the field, and strong connections to collaborators are likely to be even more challenged. Nevertheless, we will continue to advance the use of shared data to provide answers needed to improve the life of patients with kidney disease.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.548
metaresearch head score (Gemma)0.446
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch, Open science
Consensus categoriesMetaresearch
DomainCandidate signal: Reproducibility · Consensus signal: none
Study designCandidate signal: Qualitative · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.985
Threshold uncertainty score0.557

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.5480.446
Meta-epidemiology (narrow)0.0010.002
Meta-epidemiology (broad)0.0030.003
Bibliometrics0.0060.011
Science and technology studies0.0180.037
Scholarly communication0.0360.049
Open science0.0150.072
Research integrity0.0190.029
Insufficient payload (model declined to judge)0.0070.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.050
GPT teacher head0.354
Teacher spread0.304 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designQualitative
DomainReproducibility
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueJournal of the American Society of Nephrology→Same topicChronic Kidney Disease and Diabetes→French-language works237,207→