MétaCan
Menu
Back to cohort
Record W2092100939 · doi:10.1093/jnci/djn074

Should Observational Studies Be a Thing of the Past?

2008· letter· en· W2092100939 on OpenAlexaff
Kathleen I. Pritchard

Bibliographic record

VenueJNCI Journal of the National Cancer Institute · 2008
Typeletter
Languageen
FieldMedicine
TopicCancer Risks and Factors
Canadian institutionsSunnybrook Health Science Centre
Fundersnot available
KeywordsObservational studyMedicinePathology

Abstract

fetched live from OpenAlex

Holmberg et al. ( 1 ) have presented, in this issue of the Journal, updated follow-up on their randomized HABITS (Hormonal Replacement Therapy After Breast Cancer—Is it Safe?) trial of the use of hormone replacement therapy (HRT) in breast cancer survivors. This topic has been extremely controversial. Before the initial publication of the HABITS trial ( 2 ), the only available data concerning the use of HRT in breast cancer survivors came from a series of observational studies and clinical case series, the results of which suggested that this practice was safe ( 3 ). A small, underpowered randomized trial also provided some reassurance ( 4 ). Now, Holmberg et al.'s ( 1 ) follow-up of the HABITS trial suggests quite definitively that there is a statistically significantly increased risk of recurrence in women given HRT following a diagnosis of breast cancer (hazard ratio = 2.4, 95% CI = 1.3 to 4.2). Fewer than 500 randomly assigned patients were required to demonstrate a difference in the incidence of new breast cancer events at 5 years: that is, 22% in the hormone therapy arm vs 8% in the control arm of this trial. Although another study that was completed around the same time, the Stockholm Breast Cancer Study Group Trial ( 5 ), does not show as clear an effect, pooled data from the two studies still suggested a harmful effect from the use of HRT. Both the Stockholm and HABITS trials were terminated early because of these safety analyses ( 6 ). Although randomized data concerning use of HRT for symptomatic intervention in breast cancer survivors are still sparse, it seems that the harmful side effects of HRT have finally been clearly demonstrated in what is, by today's standards, a small randomized trial, carried out in a few relatively small countries. Why did we wait so long? The controversy in this area parallels a larger controversy relating to the use of HRT in healthy women. During the same time period that the HABITS trial was ongoing, a randomized trial designed to measure the overall health effects of HRT using estrogen alone in hysterectomized women and estrogen plus progesterone in women with an intact uterus was finally being carried out ( 7 ). This trial, as well, has contradicted 30 years of dogma concerning the use of HRT. Before the publication of the Women's Health Initiative (WHI) Study ( 7 ), there was controversy regarding the benefits of such intervention. In general, it was agreed that HRT reduced hot flashes, improved general well-being, and protected bone in peri- and postmenopausal women. Randomized trials have not changed our understanding of these benefits, but virtually all other dogma concerning the use of HRT in healthy women has been turned topsy-turvy. Before the WHI results, HRT, particularly estrogen, was believed to protect the heart and was widely hypothesized to improve cognition. HRT was believed to be associated with an increased risk of breast cancer, but this was thought to be similar whether progesterone was used or not. Now the WHI results have shown, with the robust methodology of a randomized trial, that estrogen alone does not clearly increase the risk of breast cancer. However, combined use of estrogen and progesterone clearly does. Furthermore, neither estrogen nor the combination of estrogen and progesterone improved cognition, and the combination increased the incidence of stroke and dementia. Coronary artery disease (CAD) was not reduced by the use of estrogen alone, and estrogen plus progesterone increased CAD. Could we have been more wrong? Unanswered questions remain regarding the role of HRT in women with a previous diagnosis of breast cancer. Might the use of estrogen alone be safer than the use of estrogen combined with a progesterone? Can topical estrogen products such as the Estring or vaginal estrogen creams be used with safety? Additional randomized trials or additional mining of data from completed randomized trials may usefully increase the amount of reliable data. However, one wonders why we took the results of observational studies as seriously as we did. Why did the observational studies so mislead us? As Holmberg et al. ( 1 ) state in their current publication, “it is not surprising that the results from this randomized trial deviate from those in the observational series.” The bias inherent in selecting breast cancer survivors for an HRT trial would seem obvious. Every clinician with a patient considering such therapy would be likely to have screened for the presence of undetected metastatic disease. In addition, women at lower risk for recurrence or women who had gone for long periods without recurrence would be more likely to be entered into such case series or observational studies. Attempts to adjust for the effects of these biases were clearly inadequate. It seems ridiculous to continue to impute effects from observational studies when the conduct of a relatively small randomized controlled clinical trial could clearly provide a definitive answer to the question under study. There are situations in the management of breast cancer that are not amenable to a randomized clinical trial. For example, we rely on observational data to “help” us to advise women who wish to become pregnant following a diagnosis of breast cancer. Basic biology would suggest that pregnancy is risky in women with hormone-responsive disease. We use observational data from women who are clearly highly self- and doctor-selected for pregnancy to suggest that pregnancy is safe or perhaps even advantageous. This approach is methodologically flawed because the selection bias in this situation is probably even greater than in observational studies of the use of HRT in breast cancer survivors. However, we can do little else than use these data, together with an explanation of their potential inadequacies, in advising women because no randomized controlled trial of pregnancy following breast cancer diagnosis will ever be carried out. In settings such as the HRT controversy, however, randomized trials such as HABITS were long overdue. It is to be hoped that we can learn from this experience to move quickly to interventional studies with robust controlled designs in settings in which observational data clearly have the potential to mislead us.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.065
metaresearch head score (Gemma)0.247
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Methods · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Commentary · Consensus signal: Commentary
Teacher disagreement score0.935
Threshold uncertainty score0.345

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0650.247
Meta-epidemiology (narrow)0.0010.002
Meta-epidemiology (broad)0.0040.002
Bibliometrics0.0020.002
Science and technology studies0.0050.009
Scholarly communication0.0080.014
Open science0.0040.003
Research integrity0.0830.083
Insufficient payload (model declined to judge)0.0070.006

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.456
GPT teacher head0.443
Teacher spread0.013 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
DomainMethods
GenreCommentary

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations14
Published2008
Admission routes1
Has abstractno

Explore more

Same venueJNCI Journal of the National Cancer InstituteSame topicCancer Risks and FactorsFrench-language works237,207