Methodological progress note: Utilizing inpatient data sources to minimize common biases when emulating target trials
Bibliographic record
Abstract
Randomized controlled trials (RCTs) provide the highest level of evidence, however it may not be feasible, ethical, or timely, to perform an RCT for all research questions. In these situations, observational data can be used to research the study question at hand. Designing an observational study to emulate a target trial can help to minimize biases present in typical observational studies. The concept of the target trial has been made popular by Hernan and Robins, and refers to the ideal randomized trial that would best answer the specific study question, were it completed.1 Target trial emulation is the process of carefully designing an observational study such that it mirrors the target trial as closely as possible on all study components. This involves first, articulating a study protocol for the target randomized trial, and second, explicitly translating each study component to be applied in an observational setting. This includes defining the eligibility criteria, treatments groups, follow-up period, outcome ascertainment, and analysis plan, as they would be designed and implemented in a randomized trial.1 Due to these requirements, target trial emulation is most appropriate when the research question could theoretically be answered by an RCT, which is often the case for cohort and longitudinal studies, but not for case–control or cross-sectional studies. Most target trial emulations to date have focused on the outpatient setting.2 However, several important biases, such as immortal time bias, time-varying confounding, and measurement bias, can be difficult to eliminate when using outpatient data sources. In contrast, inpatient data sources typically contain a robust breadth and depth of data that can be utilized to mitigate these biases through target trial emulation, and are ideal for the evaluation of interventions and outcomes that occur during a hospitalization. This article will review the most common sources of bias in target trial emulations and review how to minimize these sources of bias using inpatient data. One of the most important sources of bias in nonrandomized studies is immortal time bias.3 Immortal time occurs when there is a delay between when a patient meets eligibility criteria and is included in the study, and when follow-up to assess for outcomes begins.3 Between these two time points, any outcomes will not be counted. The amount of immortal time, during which outcomes can not occur by design, may vary between treatment groups, which can subsequently lead to immortal time bias. This can occur in the outpatient setting, where eligibility criteria may be measured at multiple time points and may differ between patients. For example, consider an outpatient cohort study evaluating mortality among patients with pancreatic cancer that compares a specialty cancer center with a wait-time of 2 months, to a local clinic with a wait-time of 1 week. If follow-up time is measured from the time of cancer diagnosis, all patients who are referred to the specialty cancer center but died before being seen would be excluded from the analysis. In effect, only patients who survived until they were seen by the specialty cancer center (at least 2 months) would be included in the analysis. In this way, even if there was no actual difference in treatment effectiveness between the two groups, the specialty cancer center group would appear to have a longer survival because of an “immortal time” of 2 months. However, when using data from the inpatient setting, the index date can be standardized between all participants (e.g., by utilizing the first day of hospitalization). In doing so, both groups of patients have the same time-zero (i.e., the time at which follow-up begins) which prevents a delay between study inclusion and outcome assessment, and therefore minimizes the possibility of immortal time bias. Additionally, eligibility criteria should be critically assessed for both groups, to ensure that there are no criteria which may systematically vary in time of occurrence (eg., procedures, investigations, or subspecialty consultations, which may not occur on evenings or weekends). A second important source of bias in observational studies is time-varying confounding.1 This is particularly important when the time between the exposure and the outcome is months or years. For example, a cohort study evaluating the risk of heart failure among patients with osteoarthritis prescribed a nonsteroidal anti-inflammatory drug (NSAID) versus an opioid for pain control may be balanced at cohort entry, but if one group develops coronary artery disease at a greater rate over the follow-up period for a reason unrelated to their use of the intervention, and subsequently develops the outcome more frequently, the results may suggest an association that is in fact due to an unmeasured, time-varying confounder (i.e., coronary artery disease). In contrast, the short outcome time-window and availability of granular day-to-day data limit the effect of time-varying confounding in inpatient target trial emulations (median length of stay in internal medicine is 4 days).4 Furthermore, because of the availability of granular data throughout the hospitalization, sensitivity analyses can be conducted to ensure that time-varying confounding has not significantly affected the study results for key variables, which is often not possible in the outpatient setting.1 In addition to time-varying confounding, time-varying exposures, such as medication adherence, can also bias study results.1, 5 In the outpatient setting, medication use and adherence is impossible to know with certainty, and instead is inferred from pharmacy prescription dispensing.2 For example, in pharmacoepidemiology studies, if a medication is re-filled 90 days after the first prescription and a total supply of 90 additional days are provided, then it is assumed the medication was taken for all 180 days. This of course is an imperfect measure of adherence. Furthermore, outpatient pharmacoepidemiology studies do not include non-prescription medications, which can be an important source of unmeasured exposures (e.g., over the counter proton pump inhibitor use).5 In contrast, within the inpatient setting, medication use is known with near certainty, as medication administration is monitored and documented, including medications that would typically be obtained (such as NSAIDs and tylenol). This availability of more complete and accurate data helps to limit bias arising from time-varying exposures in the inpatient setting. Additionally, conducting the primary analysis using an intention-to-treat approach (i.e., group at study enrollment), with sensitivity analyses for treatment cross-over and variable censoring, can help to limit the impact of time-varying exposures. Measurement bias is a common problem in outpatient observational studies that utilize health administrative or insurance data.1, 2, 5 Measurement bias can be present for both baseline patient variables as well as outcome ascertainment. Most outpatient studies identify their outcome, and confounders, using International Classification of Diseases (ICD) 10 codes. While some ICD-10 codes are accurate, many are not.6 For example the sensitivity of ICD-10 codes for iron deficiency anemia and electrolytes disorders are 30.3% and 36.3%, respectively.6 Measurement bias can be minimized using inpatient data by selecting clinically relevant outcomes which can be directly extracted from the medical record (e.g., hemoglobin/electrolyte values, hypoxemia, hypotension, or death), as opposed to relying on diagnostic codes that can be either inaccurate or unvalidated. Attrition bias can be a significant concern in observational studies conducted in outpatient populations where data is not captured for a participant during the entire follow-up window.1, 2, 5 Depending on the source of health administrative data utilized, this can occur due to individual's moving geographic location, changing insurance plans, or not attending follow-up appointments or investigations.2 Attrition bias can prevent important outcomes from being observed and recorded. For example, if two anticoagulants are being compared, and one more frequently causes an adverse effect that is not recorded in the database but leads to treatment discontinuation or not attending follow-up, these patients will be censored from the analysis and not contribute clinically important information. In contrast, because inpatient trials occur while a patient is in a monitored setting and undergoing multiple daily evaluations, most clinically important events that occur while a patient is hospitalized are known and available for analysis. Emulation of a target trial with observational data is dependent upon an accurate method of balancing baseline characteristics between treatment groups. In the target trial this would be achieved through randomization. In observational studies, various strategies, such as matching, stratification, regression, and use of propensity score techniques, can be used.1, 5, 7 These strategies can only ensure balance among the measured variables that are included. Within many outpatient cohorts, potentially important confounding variables are often not available either at all or at the time of study entry (e.g., hemoglobin A1c values in studies of patients with diabetes, blood pressure for studies of antihypertensives, etc.). In contrast, an inpatient hospitalization is a “data-rich” time window in which individuals have many characteristics and tests measured that can be utilized to optimize the balance of confounding variables to better simulate randomization (though, importantly, all baseline covariates must be measured before treatment initiation or group assignment).1 Furthermore, because of the availability of data during a hospitalization, a large number of potential “control” or “negative” outcomes can be evaluated to indirectly assess for signs that balance was achieved between the groups. Negative outcomes are evaluations between treatment groups where a difference in effect is not expected, and if confirmed, lend additional support to the groups being balanced in important baseline covariates. For example, in a cohort study comparing prasugrel to ticagrelor, a negative outcome of urinary tract infections (which are not known to be associated with either medication) could be included to indirectly evaluate group balance and potential unmeasured confounding. The benefits of well-conducted observational studies have been described elsewhere and include studying treatment effects in a real-world population and including participants who would typically be excluded from RCTs.1, 2 This often includes participants who have limited proficiency in the national language, low socioeconomic status, are experiencing homelessness, or do not have the social or financial resources to attend research study visits.8 Because most admissions are on an emergent basis, inpatient populations are more representative of the target population and, in jurisdictions where healthcare is publicly funded, are not influenced by financial resources. Furthermore, they are not dependent upon enrollment in a private insurance plan or attending outpatient appointments (Table 1). Target trial emulation using inpatient data does have limitations. First, because treatment group assignment is not random, confounding by indication or unmeasured confounding always remains a possibility. Second, inpatient care is dynamic, with frequent changes to patient conditions and treatment strategies, and as such, isolating the effect of a single intervention may be challenging if the study is not designed appropriately. Third, measurements are often conducted using the best available data, as opposed to assessments that are standardized in both timing and content, as would be conducted for a clinical trial. Fourth, inpatient target trials are best suited to assess outcomes that occur during the hospitalization (e.g., ICU transfer, inpatient mortality). Fifth, it is important to perform a sample size calculation as part of protocol development to ensure the study is adequately powered to detect the outcome of interest. Sixth, when conducting multisite target trial emulations, it is imperative that study protocol components are applied similarly at all sites (e.g., inclusion/exclusion criteria, outcome assessment, etc.). Finally, it is not the inpatient data sources in and of themselves that minimize bias, but rather that the robustness of inpatient data allows for the application of target trial emulation principles, which in turn minimizes bias. As such, if target trial emulation practices are not correctly applied, inpatient studies remain at risk of common biases. Furthermore, outpatient data sources continue to improve and in many settings are sufficient to perform high-quality target trial emulation. In summary, inpatient target trial emulations can be used to overcome many of the limitations of outpatient observational studies, including immortal time bias, measurement bias, time-varying confounding, unmeasured confounding, and attrition bias, when the principles outlined by Hernan and Robins are applied.1, 2 High quality observational studies are needed to address common and important inpatient medical problems. M. F. was a consultant for ProofDx, a start-up company focused on creating a point-of-care diagnostic test for COVID19. M. F. is also an advisor for Signal1, a start-up company that deploys machine learned models to improve patient care.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.013 | 0.024 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".