MétaCan
Menu
Back to cohort
Record W4401103214 · doi:10.1093/jbmr/zjae099

Improving the value and interpretation of observational studies comparing treatment effects of osteoporosis medications depends on standardized reporting of methods

2024· article· en· W4401103214 on OpenAlexaboutno aff
Kaleen N. Hayes, Arman Oganisian, Douglas P. Kiel

Bibliographic record

VenueJournal of Bone and Mineral Research · 2024
Typearticle
Languageen
FieldMedicine
TopicBone health and osteoporosis research
Canadian institutionsnot available
FundersNIH Office of the DirectorOffice of Disease PreventionNational Institute on AgingNational Institutes of Health
KeywordsVeterans AffairsObservational studyPublic healthMedicineInterpretation (philosophy)Medical schoolFamily medicineGerontologyHistoryLibrary scienceMedical educationNursingInternal medicine

Abstract

fetched live from OpenAlex

With the approval of multiple drugs to treat osteoporosis and a drying of the pipeline for new drug development, there has been a heightened interest in using large observational studies to compare the benefits and harms of different osteoporosis drugs. Only a handful of clinical trials have compared new medications to previously approved drugs. Large observational studies also add the dimension of “real world” experience, longer durations of follow-up, and greater sample sizes that cannot be achieved in most clinical trials. Observational studies that leverage real-world data sources can be powerful tools to conduct comparisons of osteoporosis medications, and there have been recent calls to adopt formal frameworks for their use to estimate causal effects in the medical literature.1 Propensity score (PS) methods are a common tool to attempt to mitigate bias in these studies from the non-randomized nature of treatment selection for real-world patients. Applications of PS methods can vary widely, and thus transparency in reporting of PS methods provides key information for the interpretation of findings when they are used. In this issue of the Journal of Bone and Mineral Research, such observational studies by Jeon et al. and Curtis et al. leverage PS methods to compare the effects of initial treatment with oral bisphosphonates versus denosumab.2,3 We congratulate authors on this important progress in addressing the clinically meaningful problem of equipoise between initial therapy with denosumab versus bisphosphonates. Jeon et al. leveraged national Korean healthcare data to compare outcomes among adults aged 50 yr or older newly initiating any oral bisphosphonate or denosumab for osteoporosis between January 2018 and April 2022. From the eligible population, 45 730 users of denosumab were PS-matched 1:1 to 45 730 users of oral bisphosphonate. No significant difference in relative fracture risk for denosumab vs bisphosphonates was identified in an as-treated analysis for major osteoporotic fracture (hazard ratio (HR) = 1.13 [95% CI, 0.97–1.32]), hip/pelvis fractures (HR = 1.12 [0.85–1.48]), or vertebral fractures (HR = 1.03 [95% CI, 0.81–1.31]). In subgroup analyses, those without a prior fracture did appear to have a slightly reduced relative risk of fracture with denosumab (HR = 0.91 [95% CI, 0.83–0.99]). Curtis et al. reported the comparative effectiveness of the 2 treatments in a cohort study of women aged 66 yr or older who were US Medicare Fee-for-Service beneficiaries newly initiating denosumab or alendronate between 2012 and 2018. The investigators used augmented inverse probability of treatment weighting (AIPTW), a variation of the usual inverse probability of treatment weighting (IPTW) method. IPTW adjusts for observed confounding by first estimating the conditional probability of treatment for each participant, given confounders (ie, the PS), then weighting each individual by the inverse of the probability of their observed treatment (1/PS for treated subjects and 1/[1-PS] for untreated subjects).4 This weighting balances the distribution of observed confounders between treatment groups. However, with IPTW (and any other PS method), estimates can be biased if the PS model is misspecified (eg, an interaction between covariates exists but is not included). AIPTW offers robustness to misspecification by adjusting, or “augmenting,” the usual IPTW weight using the predicted probability of the outcome based on an individual’s covariates (ie, from an “outcome model”).5 Even if the PS model is misspecified, AIPTW will still provide unbiased estimates if the outcome model is correct. This is referred to as “double-robustness” as it offers another chance to avoid bias. Curtis et al. additionally employed inverse-probability censoring weights (IPCWs).6 Just as IPTW balances the distribution of observed covariates between treatment groups, IPCW attempts to balance the distribution of covariates between censored and uncensored subjects—which corrects for informative censoring due to observed covariates (eg, if denosumab patients who are older are both more likely to be censored and have fracture events). After weighting, Curtis et al. reported using an as-treated analysis that patients initiating denosumab versus alendronate experienced a reduced risk ratio of major osteoporotic fracture (RR = 0.61 [95% CI, 0.48–0.74]), hip fracture (RR = 0.64 [95% CI, 0.39–0.90]), and other fracture outcomes. The difference in hospitalized vertebral fracture was not statistically significant, but the point estimate also suggests reduced risk (RR = 0.70 [95% CI, 0.40–1.01]). Results were consistent when stratifying analyses by fracture history. Of note, these estimates are similar to those in the original landmark trial of denosumab (eg, effect of denosumab versus placebo on hip fractures: HR = 0.60 [95% CI, 0.37–0.97]), which suggests they may be over-estimated.7 Ultimately, these studies together suggest that initial treatment with denosumab is similarly or moderately more effective than oral bisphosphonates at reducing fracture risk. However, understanding differences in results between the studies is important for clinical interpretation and future research. There are several potential explanations for the disparities in results. The first is that different populations were studied over different lengths of follow-up. The population examined in Jeon et al. was overall at lower fracture risk than those included by Curtis et al.: they were younger (46% vs 71% were ≥ 70 yr) and presumably comprised fewer White individuals (vs 83% in Curtis et al.). Furthermore, while Curtis et al. included only new initiators of alendronate in the bisphosphonate group, Jeon et al. also included new initiators of ibandronate and risedronate. Ibandronate has been shown to provide less nonvertebral fracture risk reduction than alendronate and risedronate.8 Though the distribution of bisphosphonates was unclear, ibandronate made up approximately one-third of bisphosphonate sales in Korea in 2018 and so likely represented a sizable proportion of use in Jeon et al.9 Though not reported, differences in adherence to bisphosphonate therapy, including stopping and starting patterns that are common after initiation of treatment,10 may also have contributed to disparate results. Importantly, the length of follow-up in Jeon et al. was also much shorter (median for both treatment groups = 5.6 mo). Follow-up was longer but differed between treatment groups in Curtis et al. (eg, for the hip fracture outcome, median 7.2 mo in alendronate vs 14.4 mo in the denosumab group). Risk ratios in Curtis et al. suggested greater fracture reduction with denosumab later in follow-up. One-year fracture risk estimates in Curtis et al. were numerically closer to those in Jeon et al, though still ultimately on the opposite side of the null value (eg, 1-yr RR for major osteoporotic fracture of 0.91 [95% CI, 0.85–0.97] in Curtis et al. versus HR = 1.13 [95% CI, 0.97–1.32] in Jeon et al.). Second, the results of one or both studies possibly were subject to bias from residual confounding, as neither study had information on BMD values or laboratory tests (eg, markers of bone turnover) in the primary analysis. Curtis et al. did conduct a quantitative bias analysis (QBA) for a subset of patients with linked electronic health record (EHR) data containing BMD that was used to calculate fracture risk scores; the QBA suggests that results were robust to reasonable levels of unmeasured confounding due to different baseline fracture risk and BMD. Furthermore, the primary analyses in both studies accounted for hundreds (Curtis et al.) to thousands (Jeon et al.) of potential confounding factors, and covariates were balanced between groups after PS matching or weighting. Better understanding of the exact covariates used in PS models, particularly in Jeon et al., would help to shed further light on differences in control for confounding between studies. However, different PS methods might perform differently in reducing confounding, even when the same measured covariates are used. Another article in this issue of the Journal of Bone and Mineral Research by Tan et al. directly compares the performance of PS methods using negative control outcomes (outcomes that would presumably have no relationship to the treatment except through confounding [eg, ingrown nail).11 Results suggest that all methods had evidence of residual confounding when comparing initial users of denosumab and oral bisphosphonates. However, IPTW methods resulted in the lowest estimated magnitude of residual confounding. Thus, confounding may have been more thoroughly addressed in Curtis et al. versus Jeon et al. A third factor to consider when comparing results is the difference in PS methods that the authors employed, which result in different populations in the analysis. Jeon et al. used PS matching methods, which excluded 7002 people starting denosumab who did not have a similar counterpart in the bisphosphonate group, and by nature of 1:1 matching also removed about 74% of eligible bisphosphonate users. IPTW includes the entire population, but has trade-offs. First, IPTW may potentially include persons who are not the population of interest because they have a contraindication for the other treatment. Second, assigning extreme weights to individuals during the IPTW approach can have a large impact on the results, though stabilizing and trimming weights, as conducted by Curtis et al., reduces outliers. Finally, IPTW and 1:1 PS matching in fact target different estimands (quantities). IPTW in Curtis et al. targeted an overall average treatment effect (ATE) of initial therapy with denosumab vs alendronate among the entire study population. In contrast, 1:1 PS matching by Jeon et al. targets an average effect (ie, fracture risk with denosumab vs alendronate) among those treated (ATT) with denosumab. In principle, even if all confounders are measured and no model misspecification is present, estimates may differ since they are targeting different estimands. Of course, in any given analysis, the 2 estimands may be close numerically but, in general, these 2 estimates will not be identical. The estimand that a PS method calculates (ATE vs ATT) and its interpretation are helpful to report in the context of a particular study. Part of our inability to confidently state whether these findings differ primarily due to differences in methods vs other factors is due to differences in study reporting. Certain elements of observational study reporting are universal regardless of the analytic methods used. For example, both the Strengthening the Reporting of Observational Studies in Epidemiology reporting guidelines for cohort studies12 and the Reporting of Studies Conducted using Observational Routinely Collected Health Data Statement for Pharmacoepidemiology 13 recommend all cohort studies report unadjusted estimates (ie, regression models without IPTW or before matching) and counts of outcomes and censoring events. Reporting crude estimates and counts provides context for how analytic methods to address bias modify results. Furthermore, estimating and reporting the absolute effect measures (ie, risk differences at specified time points) are recommended for all observational studies.14 Providing absolute risks alongside relative estimates (eg, HRs) is critical for clinical decision-making and to avoid over-interpretation of results.15 For example, in Curtis et al., at 2 yr of follow-up, the denosumab users had a 12% lower relative risk of major osteoporotic fracture (RR = 0.88), yet visual approximations from the weighted cumulative incidence curves of the absolute risks of fracture in each group show an absolute difference in fracture risk of about 0.5% or less. Furthermore, though PS methods are extremely common in observational studies of drug effects, standardized guidance on their reporting has not been established, particularly for osteoporosis studies. Disease-specific reporting guidance for PS methods does exist, such as those for cancer-related studies.16 Central reporting items that would help to improve interpretation in these studies include but are not limited to: (1) listing and defining all variables included in PS models and their operationalization in models (eg, binary, continuous); (2) specifying the matching method and caliper employed, when applicable; and (3) reporting the mean and distribution of both IPTW and censoring weights before and after trimming. Furthermore, any use of doubly robust or augmented methods should be described in detail, including the exact outcome models specified. We believe at least one negative control outcome should also be examined to better understand the degree to which PS methods addressed confounding.11,17 In conclusion, these studies together suggest that initial treatment with denosumab is similarly or moderately more effective than oral bisphosphonates at reducing fracture risk. Choice of osteoporosis therapy should be informed by patient preferences and access, as well as the risk of rebound fractures upon discontinuation of denosumab.18 Reporting guidelines for PS methods that could be widely applied to studies such as those of Jeon and Curtis will facilitate interpretation of observational comparative effects studies for clinical decision-making. This editorial was funded in part by R01AG078759 funded by The National Institute on Aging (NIA), the NIH Office of the Director, and the NIH Office of Disease Prevention (ODP). K.N.H. has received grant funding paid directly to Brown University for investigator-initiated research from Sanofi, Genentech, and GlaxoSmithKline for research on influenza vaccination in nursing homes, influenza outbreak control, and shingles vaccination in nursing homes, respectively. K.N.H. has also served as a consultant for Canada’s Drug Agency (formerly the Canadian Agency for Drugs and Technologies in Health) for the development of reporting guidance for real-world evidence. D.P.K. has received grant funding to his institution for a competitive RFA on osteoporosis research from Amgen related to high resolution peripheral quantitative computed tomography and fracture risk. He has received grant funding to his institution from Solara Bio on gut microbiome research. D.P.K. serves on scientific advisory boards for Radius Health and Solarea Bio, and on a data safety committee for Agnovos. A.O. has no disclosures to report.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.824
metaresearch head score (Gemma)0.935
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesMetaresearch
DomainCandidate signal: Reporting · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: none
GenreCandidate signal: Methods · Consensus signal: Methods
Teacher disagreement score0.176
Threshold uncertainty score0.217

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.8240.935
Meta-epidemiology (narrow)0.0060.005
Meta-epidemiology (broad)0.0090.009
Bibliometrics0.0190.022
Science and technology studies0.0040.017
Scholarly communication0.0160.016
Open science0.0140.015
Research integrity0.0110.021
Insufficient payload (model declined to judge)0.0060.003

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.224
GPT teacher head0.536
Teacher spread0.312 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; the direct Gemma label and the distilled Codex classifier agree on what is shown here.

Study designNot applicable
DomainReporting
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2024
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Bone and Mineral ResearchSame topicBone health and osteoporosis researchFrench-language works237,207