Improving the value and interpretation of observational studies comparing treatment effects of osteoporosis medications depends on standardized reporting of methods
Notice bibliographique
Résumé
With the approval of multiple drugs to treat osteoporosis and a drying of the pipeline for new drug development, there has been a heightened interest in using large observational studies to compare the benefits and harms of different osteoporosis drugs. Only a handful of clinical trials have compared new medications to previously approved drugs. Large observational studies also add the dimension of “real world” experience, longer durations of follow-up, and greater sample sizes that cannot be achieved in most clinical trials. Observational studies that leverage real-world data sources can be powerful tools to conduct comparisons of osteoporosis medications, and there have been recent calls to adopt formal frameworks for their use to estimate causal effects in the medical literature.1 Propensity score (PS) methods are a common tool to attempt to mitigate bias in these studies from the non-randomized nature of treatment selection for real-world patients. Applications of PS methods can vary widely, and thus transparency in reporting of PS methods provides key information for the interpretation of findings when they are used. In this issue of the Journal of Bone and Mineral Research, such observational studies by Jeon et al. and Curtis et al. leverage PS methods to compare the effects of initial treatment with oral bisphosphonates versus denosumab.2,3 We congratulate authors on this important progress in addressing the clinically meaningful problem of equipoise between initial therapy with denosumab versus bisphosphonates. Jeon et al. leveraged national Korean healthcare data to compare outcomes among adults aged 50 yr or older newly initiating any oral bisphosphonate or denosumab for osteoporosis between January 2018 and April 2022. From the eligible population, 45 730 users of denosumab were PS-matched 1:1 to 45 730 users of oral bisphosphonate. No significant difference in relative fracture risk for denosumab vs bisphosphonates was identified in an as-treated analysis for major osteoporotic fracture (hazard ratio (HR) = 1.13 [95% CI, 0.97–1.32]), hip/pelvis fractures (HR = 1.12 [0.85–1.48]), or vertebral fractures (HR = 1.03 [95% CI, 0.81–1.31]). In subgroup analyses, those without a prior fracture did appear to have a slightly reduced relative risk of fracture with denosumab (HR = 0.91 [95% CI, 0.83–0.99]). Curtis et al. reported the comparative effectiveness of the 2 treatments in a cohort study of women aged 66 yr or older who were US Medicare Fee-for-Service beneficiaries newly initiating denosumab or alendronate between 2012 and 2018. The investigators used augmented inverse probability of treatment weighting (AIPTW), a variation of the usual inverse probability of treatment weighting (IPTW) method. IPTW adjusts for observed confounding by first estimating the conditional probability of treatment for each participant, given confounders (ie, the PS), then weighting each individual by the inverse of the probability of their observed treatment (1/PS for treated subjects and 1/[1-PS] for untreated subjects).4 This weighting balances the distribution of observed confounders between treatment groups. However, with IPTW (and any other PS method), estimates can be biased if the PS model is misspecified (eg, an interaction between covariates exists but is not included). AIPTW offers robustness to misspecification by adjusting, or “augmenting,” the usual IPTW weight using the predicted probability of the outcome based on an individual’s covariates (ie, from an “outcome model”).5 Even if the PS model is misspecified, AIPTW will still provide unbiased estimates if the outcome model is correct. This is referred to as “double-robustness” as it offers another chance to avoid bias. Curtis et al. additionally employed inverse-probability censoring weights (IPCWs).6 Just as IPTW balances the distribution of observed covariates between treatment groups, IPCW attempts to balance the distribution of covariates between censored and uncensored subjects—which corrects for informative censoring due to observed covariates (eg, if denosumab patients who are older are both more likely to be censored and have fracture events). After weighting, Curtis et al. reported using an as-treated analysis that patients initiating denosumab versus alendronate experienced a reduced risk ratio of major osteoporotic fracture (RR = 0.61 [95% CI, 0.48–0.74]), hip fracture (RR = 0.64 [95% CI, 0.39–0.90]), and other fracture outcomes. The difference in hospitalized vertebral fracture was not statistically significant, but the point estimate also suggests reduced risk (RR = 0.70 [95% CI, 0.40–1.01]). Results were consistent when stratifying analyses by fracture history. Of note, these estimates are similar to those in the original landmark trial of denosumab (eg, effect of denosumab versus placebo on hip fractures: HR = 0.60 [95% CI, 0.37–0.97]), which suggests they may be over-estimated.7 Ultimately, these studies together suggest that initial treatment with denosumab is similarly or moderately more effective than oral bisphosphonates at reducing fracture risk. However, understanding differences in results between the studies is important for clinical interpretation and future research. There are several potential explanations for the disparities in results. The first is that different populations were studied over different lengths of follow-up. The population examined in Jeon et al. was overall at lower fracture risk than those included by Curtis et al.: they were younger (46% vs 71% were ≥ 70 yr) and presumably comprised fewer White individuals (vs 83% in Curtis et al.). Furthermore, while Curtis et al. included only new initiators of alendronate in the bisphosphonate group, Jeon et al. also included new initiators of ibandronate and risedronate. Ibandronate has been shown to provide less nonvertebral fracture risk reduction than alendronate and risedronate.8 Though the distribution of bisphosphonates was unclear, ibandronate made up approximately one-third of bisphosphonate sales in Korea in 2018 and so likely represented a sizable proportion of use in Jeon et al.9 Though not reported, differences in adherence to bisphosphonate therapy, including stopping and starting patterns that are common after initiation of treatment,10 may also have contributed to disparate results. Importantly, the length of follow-up in Jeon et al. was also much shorter (median for both treatment groups = 5.6 mo). Follow-up was longer but differed between treatment groups in Curtis et al. (eg, for the hip fracture outcome, median 7.2 mo in alendronate vs 14.4 mo in the denosumab group). Risk ratios in Curtis et al. suggested greater fracture reduction with denosumab later in follow-up. One-year fracture risk estimates in Curtis et al. were numerically closer to those in Jeon et al, though still ultimately on the opposite side of the null value (eg, 1-yr RR for major osteoporotic fracture of 0.91 [95% CI, 0.85–0.97] in Curtis et al. versus HR = 1.13 [95% CI, 0.97–1.32] in Jeon et al.). Second, the results of one or both studies possibly were subject to bias from residual confounding, as neither study had information on BMD values or laboratory tests (eg, markers of bone turnover) in the primary analysis. Curtis et al. did conduct a quantitative bias analysis (QBA) for a subset of patients with linked electronic health record (EHR) data containing BMD that was used to calculate fracture risk scores; the QBA suggests that results were robust to reasonable levels of unmeasured confounding due to different baseline fracture risk and BMD. Furthermore, the primary analyses in both studies accounted for hundreds (Curtis et al.) to thousands (Jeon et al.) of potential confounding factors, and covariates were balanced between groups after PS matching or weighting. Better understanding of the exact covariates used in PS models, particularly in Jeon et al., would help to shed further light on differences in control for confounding between studies. However, different PS methods might perform differently in reducing confounding, even when the same measured covariates are used. Another article in this issue of the Journal of Bone and Mineral Research by Tan et al. directly compares the performance of PS methods using negative control outcomes (outcomes that would presumably have no relationship to the treatment except through confounding [eg, ingrown nail).11 Results suggest that all methods had evidence of residual confounding when comparing initial users of denosumab and oral bisphosphonates. However, IPTW methods resulted in the lowest estimated magnitude of residual confounding. Thus, confounding may have been more thoroughly addressed in Curtis et al. versus Jeon et al. A third factor to consider when comparing results is the difference in PS methods that the authors employed, which result in different populations in the analysis. Jeon et al. used PS matching methods, which excluded 7002 people starting denosumab who did not have a similar counterpart in the bisphosphonate group, and by nature of 1:1 matching also removed about 74% of eligible bisphosphonate users. IPTW includes the entire population, but has trade-offs. First, IPTW may potentially include persons who are not the population of interest because they have a contraindication for the other treatment. Second, assigning extreme weights to individuals during the IPTW approach can have a large impact on the results, though stabilizing and trimming weights, as conducted by Curtis et al., reduces outliers. Finally, IPTW and 1:1 PS matching in fact target different estimands (quantities). IPTW in Curtis et al. targeted an overall average treatment effect (ATE) of initial therapy with denosumab vs alendronate among the entire study population. In contrast, 1:1 PS matching by Jeon et al. targets an average effect (ie, fracture risk with denosumab vs alendronate) among those treated (ATT) with denosumab. In principle, even if all confounders are measured and no model misspecification is present, estimates may differ since they are targeting different estimands. Of course, in any given analysis, the 2 estimands may be close numerically but, in general, these 2 estimates will not be identical. The estimand that a PS method calculates (ATE vs ATT) and its interpretation are helpful to report in the context of a particular study. Part of our inability to confidently state whether these findings differ primarily due to differences in methods vs other factors is due to differences in study reporting. Certain elements of observational study reporting are universal regardless of the analytic methods used. For example, both the Strengthening the Reporting of Observational Studies in Epidemiology reporting guidelines for cohort studies12 and the Reporting of Studies Conducted using Observational Routinely Collected Health Data Statement for Pharmacoepidemiology 13 recommend all cohort studies report unadjusted estimates (ie, regression models without IPTW or before matching) and counts of outcomes and censoring events. Reporting crude estimates and counts provides context for how analytic methods to address bias modify results. Furthermore, estimating and reporting the absolute effect measures (ie, risk differences at specified time points) are recommended for all observational studies.14 Providing absolute risks alongside relative estimates (eg, HRs) is critical for clinical decision-making and to avoid over-interpretation of results.15 For example, in Curtis et al., at 2 yr of follow-up, the denosumab users had a 12% lower relative risk of major osteoporotic fracture (RR = 0.88), yet visual approximations from the weighted cumulative incidence curves of the absolute risks of fracture in each group show an absolute difference in fracture risk of about 0.5% or less. Furthermore, though PS methods are extremely common in observational studies of drug effects, standardized guidance on their reporting has not been established, particularly for osteoporosis studies. Disease-specific reporting guidance for PS methods does exist, such as those for cancer-related studies.16 Central reporting items that would help to improve interpretation in these studies include but are not limited to: (1) listing and defining all variables included in PS models and their operationalization in models (eg, binary, continuous); (2) specifying the matching method and caliper employed, when applicable; and (3) reporting the mean and distribution of both IPTW and censoring weights before and after trimming. Furthermore, any use of doubly robust or augmented methods should be described in detail, including the exact outcome models specified. We believe at least one negative control outcome should also be examined to better understand the degree to which PS methods addressed confounding.11,17 In conclusion, these studies together suggest that initial treatment with denosumab is similarly or moderately more effective than oral bisphosphonates at reducing fracture risk. Choice of osteoporosis therapy should be informed by patient preferences and access, as well as the risk of rebound fractures upon discontinuation of denosumab.18 Reporting guidelines for PS methods that could be widely applied to studies such as those of Jeon and Curtis will facilitate interpretation of observational comparative effects studies for clinical decision-making. This editorial was funded in part by R01AG078759 funded by The National Institute on Aging (NIA), the NIH Office of the Director, and the NIH Office of Disease Prevention (ODP). K.N.H. has received grant funding paid directly to Brown University for investigator-initiated research from Sanofi, Genentech, and GlaxoSmithKline for research on influenza vaccination in nursing homes, influenza outbreak control, and shingles vaccination in nursing homes, respectively. K.N.H. has also served as a consultant for Canada’s Drug Agency (formerly the Canadian Agency for Drugs and Technologies in Health) for the development of reporting guidance for real-world evidence. D.P.K. has received grant funding to his institution for a competitive RFA on osteoporosis research from Amgen related to high resolution peripheral quantitative computed tomography and fracture risk. He has received grant funding to his institution from Solara Bio on gut microbiome research. D.P.K. serves on scientific advisory boards for Radius Health and Solarea Bio, and on a data safety committee for Agnovos. A.O. has no disclosures to report.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,824 | 0,935 |
| Méta-épidémiologie (sens strict) | 0,006 | 0,005 |
| Méta-épidémiologie (sens large) | 0,009 | 0,009 |
| Bibliométrie | 0,019 | 0,022 |
| Études des sciences et des technologies | 0,004 | 0,017 |
| Communication savante | 0,016 | 0,016 |
| Science ouverte | 0,014 | 0,015 |
| Intégrité de la recherche | 0,011 | 0,021 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,006 | 0,003 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; l’étiquette directe de Gemma et le classifieur distillé Codex s’accordent sur ce qui est montré ici.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».