Randomised evaluation of government health programmes does present a challenge to standard research ethics frameworks
Bibliographic record
Abstract
In a recent issue of Journal of Medical Ethics (JME), we discussed the ethical review of evaluations of interventions that would occur whether or not the evaluation was taking place. We concluded that standard research ethics frameworks including the Ottawa Statement, which requires justification for all aspects of an intervention and its roll-out, were a poor guide in this area. We proposed that a consideration of researcher responsibility, based on the consequences of the research taking place, would be a more appropriate way delineate the scope of research ethics review. Weijer and Taljaard present a counterargument to our proposal, which we address in this reply. They claim that a focus on researcher responsibility will weaken the protection of research participants and link it to 'unethical research' and a 'government experimenting on its own people'. However, the moral responsibility of researchers is defined in terms of the consequences of the research on human welfare and harm, not in opposition to it. Weijer and Taljaard argue that researchers must justify what they are studying whether or not they have any control over it and that governments must justify their programmes, including by demonstrating equipoise, to a research ethics committee if they implement them in a randomised way. We strongly disagree that this is a defensible way to define the scope of research ethics review and argue that this provides no further protections to research participants beyond what we propose, but places a potential barrier to learning from government programmes.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.201 | 0.065 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.007 | 0.044 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".