Bibliographic record
Abstract
The whole aim of practical politics is to keep the populace alarmed (and hence clamorous to be led to safety) by menacing it with an endless series of hobgoblins, most of them imaginary. - H. L. Mencken[dagger] CONTEXT It is a common ritual among today's academics to submit research proposals to a group of colleagues, the Institutional Review Board (IRB; REB in Canada). When IRBs were first created in 1974,1 they were directed to assess whether a researcher's proposed project would expose the public to greater than everyday Over the last thirty years, IRBs have evolved to review research proposals against criteria well beyond the scope of their original mandate or their ostensible purpose.2 For example, instead of assessing whether proposed research would expose the public to greater than everyday IRBs now ask whether proposed research would expose members of the public to even risk. In practice minimal risk is a euphemism for zero risk, which is an impossible objective to achieve.3 Likewise, IRBs have expanded the review process beyond the issue of publie safety to pursue more nebulous agendas, including whether the proposed research is worthwhile. They now weigh expected social benefits, legal liability, and similar issues, presuming to judge at the outset that which can only be determined by examining the results. The broadening scope of IRB inquiry can charitably be described as creep,4 and because this expansion has happened in small steps over thirty years, the successive impositions were seldom challenged. However, the cumulative effect is striking,5 and there is no sign this mission creep has been stalled. Such subterfuge is typical of the way the research ethics enterprise has expanded over the years, always with the result that control of inquiry is increased with no documented evidence of enhanced subject safety. Today, particularly in the field of non-medical research, the institutional review process is more accurately described as censorship than safety screening.6 My intent here is to (1) describe how various distortions are used to defend and justify the ethics reviews, and (2) highlight some costs of the ethics enterprise that are routinely ignored. I will focus on social science and humanities research, in part because of my interests, but also because these disciplines seem most vulnerable to unwarranted censorship. When all of the results, intended and otherwise, are considered, it is clear that the constraints imposed on academic inquiry have not been accompanied by an increase in public benefits. I. BENEFITS The benefits of IRBs can be divided into two sets. The first set includes benefits IRB supporters claim accrue to the public despite the lack of reliable evidence that such benefits have materialized. My attention to these will be mainly to critique the shortcomings of the claim that procedure X improves public safety. The second set will concern benefits that accrue exclusively to regulators. These benefits are far more obvious, though they remain unacknowledged by IRB supporters. A. What is the Evidence, and How Do We Evaluate It Before I analyze the claimed benefits and the actual benefits to regulators, it is important to establish a clear understanding of what constitutes reliable evidence that can support a claim that IREs produce certain benefits. By definition, identifying something as a requires some variation of a pre-post assessment. That is, one needs the identification of a null or undesirable state to begin with, a manipulation, and then a secondary assessment that documents improvement. The former shows evidence of a need, and the latter shows evidence of the effectiveness of the manipulation, in this case the ethics review. The general claimed benefit is improved public safety, and I think the burden of proof for that is on the claimant. What is the evidence and what is its validity? …
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.001 | 0.003 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".