Comparative effectiveness research using registries, databases, and networks in women's and children's health: Time to embrace the future?
Bibliographic record
Abstract
Traditional study designs that identify exposure-outcome relations include nonexperimental designs such as case-control studies, cross sectional studies, and cohort studies as well as experimental study designs in the form of randomized controlled trials (RCTs). Both nonexperimental and experimental study designs are associated with issues of internal and external validity, time, and expenses. For example, nonexperimental study design has the limitations of nonrandom inclusion of participants leading to selection bias and an inability to control for confounders and mediators. On the other hand, experimental design is associated with a need for large infrastructure, as well as large sample sizes that consume significant time and resources and sometimes have limited external validity due to strict eligibility criteria and participation bias. Despite these limitations, these study designs have formed the basis of evaluating causes and consequences of diseases and effectiveness and safety of medical and surgical interventions, and will continue to do so for the foreseeable future. However, the limitations imposed by the inability to recruit a large enough sample size for rare conditions or conditions where disease patterns vary on the basis of individual characteristics such as age group, ethnicity, or other characteristics, as well as the financial challenges faced by research funding agencies, have led to the testing of alternative evaluation methods for different models of research design. The time for such endeavors appears to be ripe as large amounts of data are now collected in an electronic format at the patient's bedside or as part of national registries. Furthermore, data are also collected as part of networks where units with similar characteristics or geographic location have formed a practice-based clinical research network. We would like to highlight the potential advantages and practical opportunities to generate robust evidence arising from such comparative effectiveness research (CER) studies. Until now, the medical community has used data generated from large registries or networks to understand patterns of disease, epidemiology, and identify trends. This process is considered to be the use of real world data (RWD) to produce real world evidence (RWE). In order to obtain information on the effectiveness of any intervention, we still rely heavily on rigorously conducted RCTs. However, conducting RCTs on every condition and treatment is impractical and economically prohibitive. Thus, increasingly more, decision-making authorities are relying on evidence produced from the use of RWD to guide clinical care or enhance their work on policy development to deliver appropriate healthcare. Various medical specialties and subspecialties have grappled with this change in research design at different paces. Oncology groups pioneered the CER approach decades ago, a method in which two treatment or screening regimens were evaluated in a quasi-cluster randomized design.1 Gynecology, obstetrics,2 and neonatology3 are beginning to get on board with such conceptual options. Major funding organizations such as the Patient Centered Outcomes Research Institute in the US, have embraced the CER concept and funded a large registry to answer questions regarding treatment options for uterine fibroid.2 This will allow both understanding of the comparative safety of different treatments and the ability to incorporate stakeholders in the decision-making process of developing guidelines. Similarly, a clinical trial in neonatology is also planned as a registry-based pragmatic clinical trial,3 another form of CER study design. Critics of CER studies using registries, databases, and networks, cite problems with non-robust data collection leading to missing data points, some reservations regarding internal validity, unbalanced confounding, data-dredging, and non-transparency of evaluation and reporting. However, the International Society for Pharmacoeconomics and Outcomes Research (ISPOR) and the International Society for Pharmacoepidemiology (ISPE) task force suggested a baseline framework for conducting CER studies using RWD.4 They identified two main design frameworks: exploratory treatment effectiveness studies and hypothesis evaluating treatment effectiveness studies (HETE). The former lacks hypotheses but may contain enriched quality data to generate hypotheses. HETE studies evaluate intervention effect with a predefined hypothesis. When a HETE study is designed as a clinical trial that uses RWD from routine data collection, it is considered a pragmatic clinical trial (PCT)5 and has several advantages, including the participation of multiple centers, significant reduction in the need for infrastructure, the inclusion of all the patients rather than a select population6 (a common issue in traditional RCTs because of consent and capacity issues), and bringing a consistency of patient management at the unit level. The key elements of CER study design are the following: (a) a direct comparison of active treatments, (b) a typical population that is affected by the tested treatment decisions on a daily basis, and (c) a focus on using evidence to inform future care.7 The existing data collection process usually acts as a central conduit for recruitment and subsequently for data acquisition. However, ISPOR-ISPE has recommended a few suggestions to increase the confidence of end users in the study results. These included: pre-identification of the study as exploratory or HETE; protocol registration; publication of results with information on any deviation from the protocol; development of a data repository for result replication, and confirmation by others whenever possible; exploration of result reproducibility in different datasets; public reporting; and engagement of stakeholders in design, execution, analyses, and interpretation of result.4 The development of study design guidelines is an important step, and the significant stakeholder engagement in protocol development will be particularly important in easing the knowledge translation of evidence generated to implementation without significant delays. The concerns of missing data also need to be handled transparently and critically using proper statistical techniques such as imputation and accompanying sensitivity analyses. One key piece in generating robust evidence from CER using RWD is the ability to conduct sensitivity analyses, as unmeasured confounding is a significant challenge. Solutions in addition to traditional regression techniques include stratified analyses, propensity-score based analyses, instrumental variable approach, and structural modeling techniques, which are suggested to enhance confidence in results.8 If such analyses provide disparate results, then the opportunity to explore the reasons behind the discrepancy may lead to an understanding of predictors of treatment effectiveness or ineffectiveness. Even after completion of an RCT questions always remain about what predicts treatment effectiveness in some patients over others, and such an interrogation requires a large dataset. In contrast, evaluation of treatment effectiveness predictors lands perfectly in the realm of CER using a database, networks, or registries as a solid backbone. Finally, the issue of research fraud cannot be ignored. Transparency in publishing CER protocols, making the data available for secondary interrogation, and public disclosure of results may help to alleviate such concerns. The relative ease of data collection using electronic repositories, widespread interest in data mining to facilitate precision medicine, cost-efficiency, ability to study rare diseases, and the ability to learn variations in therapeutic responsiveness across widespread populations are attractive features of CER using registries, networks, or databases. Several practice-based research networks, repositories and registries exist for women's and children's health. We think it is time for researchers, stakeholders, patients, policymakers, and the public to engage in CER to explore what we do, and how we do it, with the goal of learning what matters and for whom. The information gathered may generate more hypotheses than answers, but at least we will have advanced one step in the right direction. The authors have no conflict of interest to disclose.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.598 | 0.644 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.012 | 0.007 |
| Bibliometrics | 0.014 | 0.023 |
| Science and technology studies | 0.002 | 0.013 |
| Scholarly communication | 0.022 | 0.050 |
| Open science | 0.010 | 0.011 |
| Research integrity | 0.015 | 0.013 |
| Insufficient payload (model declined to judge) | 0.008 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".