Visualizing External Validity: Graphical Displays to Inform the Extension of Treatment Effects from Trials to Clinical Practice
Bibliographic record
Abstract
BACKGROUND: In the presence of effect measure modification, estimates of treatment effects from randomized controlled trials may not be valid in clinical practice settings. The development and application of quantitative approaches for extending treatment effects from trials to clinical practice settings is an active area of research. METHODS: In this article, we provide researchers with a practical roadmap and four visualizations to assist in variable selection for models to extend treatment effects observed in trials to clinical practice settings and to assess model specification and performance. We apply this roadmap and visualizations to an example extending the effects of adjuvant chemotherapy (5-fluorouracil vs. plus oxaliplatin) for colon cancer from a trial population to a population of individuals treated in community oncology practices in the United States. RESULTS: The first visualization screens for potential effect measure modifiers to include in models extending trial treatment effects to clinical practice populations. The second visualization displays a measure of covariate overlap between the clinical practice populations and the trial population. The third and fourth visualizations highlight considerations for model specification and influential observations. The conceptual roadmap describes how the output from the visualizations helps interrogate the assumptions required to extend treatment effects from trials to target populations. CONCLUSIONS: The roadmap and visualizations can inform practical decisions required for quantitatively extending treatment effects from trials to clinical practice settings.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.020 | 0.319 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".