Clinical trials: what are we afraid of, what should we do?
Bibliographic record
Abstract
We would like to respond to an editorial in the September issue of the journal because it rehearses much confusion and many misconceptions about clinical trials.1 More importantly, the authors’ recommendations are wrong-headed and can only harm our patients and, secondarily, our specialty. The authors discuss some of the difficulties with recent randomized controlled trials (RCTs) that have not demonstrated the good outcomes our interventions were purported to deliver. The title of the editorial, that RCTs can be a ‘double-edged sword’, seems to warn the reader against something, but against what exactly? 1. Could it be that we are designing and participating in too many trials? In fact, we are collectively responsible for our field delivering poor (if any) evidence regarding the merits of our daily interventions. We need more trials, preferably trials designed and conducted by neurointerventionists. We must regain control of how to evaluate the merits of our own practice. Most importantly, if we are to offer patients care that they can trust, we must be constantly working to validate our still unvalidated interventions. What is the best way for us to do this? A trial, of course, but not just any type of trial. We will get back to this point. 2. Should we distrust the disappointing results of recent RCTs because they have ‘limitations’, as the citation from Concato1 suggests? What are we to do about trials that have design shortcomings? Should we stubbornly practice interventions that have now been shown to be harmful, albeit in trials ‘with limitations’, claiming them to be standard of care, just like an intervention that has been proven beneficial? Of course not; disappointing trial results simply mean that such interventions should only be offered within the context of better designed trials. 3. Should the readers of the editorial be warned against …
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.003 | 0.008 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".