MétaCan
Menu
Back to cohort
Record W2133000379 · doi:10.1200/jco.2012.47.6952

Evaluation of Treatment Benefit in <i>Journal of Clinical Oncology</i>

2013· editorial· en· W2133000379 on OpenAlexaff
Pamela J. Goodwin, Karla V. Ballman, Eric J. Small, Stephen A. Cannistra

Bibliographic record

VenueJournal of Clinical Oncology · 2013
Typeeditorial
Languageen
FieldEconomics, Econometrics and Finance
TopicHealth Systems, Economic Evaluations, Quality of Life
Canadian institutionsLunenfeld-Tanenbaum Research InstituteUniversity of TorontoMount Sinai Hospital
Fundersnot available
KeywordsMedicineClinical OncologyOncologyInternal medicineMedical physicsCancer

Abstract

fetched live from OpenAlex

Journal of Clinical Oncology is well-recognized for publishing manuscripts that describe treatment benefits or toxicities that have the potential to influence patient care. As such, it is critical to the editors and readers of our journal that such manuscripts present a high level of evidence that is not subject to bias or overstatement. In that regard, we have observed an increase in manuscripts that use observational study designs to assess benefit, partly due to the burgeoning interest in comparative effectiveness research (CER) coupled with easier access to large clinical and administrative databases that contain patient, treatment, and outcome data. One common example involves manuscripts describing the use of registry (eg, Surveillance, Epidemiology and End Results) and other administrative or clinical databases to analyze clinical outcomes of patients receiving different treatments. Although these studies use real world data that are often more representative of the variety of patients seen in clinical practice than is the case in randomized clinical trials (RCTs), there is a higher potential for bias and confounding with these designs, in part because treatment allocation is not randomized. Additional observational designs include time-trend studies which describe how treatment and outcomes have changed over time and attribute improvements in outcomes to more recent treatment approaches. However, because presentation, staging, concomitant care and other factors may also change over time, it is difficult to attribute these improved outcomes to any single factor, including recent treatment approaches. Likewise, modeling studies that use decision analytic or other approaches to quantify treatment benefits and harms or identify optimal treatment strategies in specific scenarios can be difficult to interpret with confidence. Many of the estimates and assumptions used in these models may not be valid, and models covering all contingencies may not be considered. As a result, the conclusions may not be correct. Thus, because of the potential for bias and confounding in observational studies, and because of potential inaccuracies in the assumptions and design features incorporated into modeling studies (despite the use of sensitivity analyses to examine the impact of variability in these estimates), these study designs are often sub-optimal to definitively demonstrate treatment benefits and harms. Given the potential limitations of these designs, in this editorial we would like to explain how we prioritize manuscripts submitted to JCO that claim to show a treatment benefit. The terminology used to describe treatment benefit can be confusing, but it is useful to make a distinction between two metrics, namely efficacy and effectiveness. By efficacy we refer to the outcome of a given treatment when administered under ideal circumstances (eg, in a defined population, with full compliance, delivered by competent physicians in a controlled environment, in the absence of comorbidity); in other words, whether an intervention works (or not) in a controlled situation. By effectiveness we refer to the outcome of a given treatment when administered in a more pragmatic (or real world) fashion, recognizing that compliance may be less than optimal, treatment settings may be diverse, expertise of care givers may vary and comorbidity may impact treatment outcomes. Typically, both efficacy and effectiveness are initially established in RCTs, which remain our gold standard for assessing treatment benefit. Meta-analyses that combine results of multiple RCTs can, at times, be useful to identify small(er) treatment effects that were not significant in individual trials but are clinically important, to examine overall treatment benefits when results of individual RCTs are conflicting, to explore patterns of treatment effects (eg, over time, in patient subsets) and to quantify rare toxicities. At JCO, meta-analyses that combine data at a patient level are prioritized over those that combine data at a study level, as they facilitate investigation of (and/or adjustment for) individual patient factors, and allow harmonization of analytic approaches and outcomes across studies. CER deserves special mention. We view research that investigates efficacy, effectiveness, and comparative effectiveness as a continuum, providing different but complementary information about treatment benefits and harms. CER that uses a randomized design is typically considered a form of effectiveness research and is evaluated at JCO in the same way as other RCTs. However, many CER studies use observational designs; they can be valuable to investigate patterns of harms and benefits of treatments in a variety of real world clinical settings, but they are susceptible to bias and confounding and do not reach the level of rigor associated with RCTs. For observational CER, JCO adopts the working definition put forward by the Institute of Medicine Committee: “CER is the generation and synthesis of evidence that compares the benefits and harms of alternative methods to prevent, diagnose, treat and monitor a clinical condition, or to improve the delivery of care. The purpose of CER is to assist consumers, clinicians, purchasers, and policy makers to make informed decisions that will improve health care at both the individual and population levels.” Key elements JOURNAL OF CLINICAL ONCOLOGY E D I T O R I A L VOLUME 31 NUMBER 9 MARCH 2

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.107
metaresearch head score (Gemma)0.477
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesMetaresearch
Consensus categoriesnone
DomainCandidate signal: Evaluation · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Editorial · Consensus signal: none
Teacher disagreement score0.893
Threshold uncertainty score0.566

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.1070.477
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0030.004
Bibliometrics0.0080.008
Science and technology studies0.0010.003
Scholarly communication0.0160.006
Open science0.0020.003
Research integrity0.0040.005
Insufficient payload (model declined to judge)0.0260.006

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.748
GPT teacher head0.644
Teacher spread0.104 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

Study designNot applicable
DomainEvaluation
GenreEditorial

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations26
Published2013
Admission routes1
Has abstractyes

Explore more

Same venueJournal of Clinical OncologySame topicHealth Systems, Economic Evaluations, Quality of LifeFrench-language works237,207