A tale of three labels: translating the JUPITER trial data into regulatory claims
Bibliographic record
Abstract
BACKGROUND: Whether a pivotal randomized trial will be interpreted in a similar and consistent manner by different regulatory agencies is uncertain as policy perspectives may play a role in data interpretation and the translation of trial results into clinical practice. PURPOSE: Using a contemporary example, to compare and contrast regulatory claims in the United States, Europe, and Canada that derive from a pivotal clinical trial. METHODS: The recently completed JUPITER trial of rosuvastatin as compared to placebo conducted among 17,802 men and women with LDL-C <130 mg/dL and hsCRP ≥ 2 mg/L, provides the only available data on rosuvastatin for primary prevention of cardiovascular disease. Thus, the JUPITER trial provides an opportunity to compare and contrast how regulatory agencies in the United States, Canada, and Europe chose to interpret an identical database. Labeling indications based on earlier statin trials of primary and secondary prevention were also reviewed. RESULTS: JUPITER demonstrated a 44% reduction (p < 0.000001) in the trial pre-specified primary endpoint (nonfatal myocardial infarction, nonfatal stroke, hospitalization for unstable angina, coronary revascularization, or cardiovascular death) with no evidence of heterogeneity across geographic regions. In response to these data, the US Food and Drug Administration label for rosuvastatin in primary prevention came closest to the actual JUPITER trial population by stipulating that those eligible for treatment should be older men and women with hsCRP >2 mg/L, plus one additional risk factor for heart disease. The Canadian label is silent on age and hsCRP (the major trial inclusion criterion), stipulating instead that treatment can be considered for those with 'at least two conventional risk factors for cardiovascular disease,' a group more inclusive than that studied. In contrast, the European Medicines Agency label limits treatment only to 'high risk individuals' ignoring hsCRP and using instead a post hoc definition of 'high risk' that comprised a subgroup of less than 10% of the study population who contributed but 67 events to the study total and did not show statistical significance when compared to placebo. None of the regulatory labels included the trial primary endpoint; instead, each focused on separate and different components of the primary endpoint. Similar discrepancies were found between European and North American regulatory agencies with regard to earlier pivotal trials of statins for primary prevention, but not for secondary prevention. LIMITATIONS: The JUPITER experience represents a case study and is not a systematic review of the regulatory decision process. CONCLUSIONS: Labeling indications can vary widely in different regulatory environments even when based on the same trial data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.040 | 0.008 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".