How pre-publication journal peer review (re)produces ignorance at scientific and medical journals: a case study
Bibliographic record
Abstract
The main goal of this paper is to explore how journal peer review produces and reproduces ignorance at scientific and medical journals. I focus on the case of pre- publication journal peer review (traditional peer review). Scientific ignorance is non- pejorative as the limits and borders of knowledge where new scientific ideas can contain new ignorance that pushes the boundaries of knowledge. Traditional peer review is an example of a ‘boundary judgement’ social form where content refers to decisions from the judgement of scientific written texts held to account to an overarching knowledge system – creating boundaries between what is and what is not considered science. Moreover, boundary judgement forms interact with the social form of scientific exchange where scientists communicate knowledge and ignorance. I investigate traditional peer review’s structural properties – elements that contribute to shaping relations in a form – to understand ignorance (re)production. Analysis of twenty-five cases with empirical and self- and third party accounts data, and data from eleven semi-structured interviews helps construct theoretical insights into how traditional peer review mostly contributes to ignorance reproduction. Reproduction owes to four structural properties: (1) contingency traditional peer review places on scientific exchange; (2) secrecy for original manuscripts and editorial judgements and decisions; (3) a relation of accountability to empiricism for editorial readers that helps construct a boundary for manuscripts, deemed as scientific or not; and (4) a relation of accountability to readers enhanced by a criterion of originality that appears to construct another boundary for manuscripts, deemed as newsworthy or not. I conclude with implications from this work set against Kuhn’s theory of paradigms. I also look to implications for authors, policymakers, editors, and journal publishers.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | Metaresearch Domain: Evaluation · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Qualitative | high |
| gpt | MetaresearchScience and technology studiesScholarly communication Domain: Evaluation · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Qualitative | high |
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.334 | 0.559 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.049 | 0.052 |
| Science and technology studies | 0.004 | 0.001 |
| Scholarly communication | 0.105 | 0.002 |
| Open science | 0.009 | 0.011 |
| Research integrity | 0.000 | 0.003 |
| Insufficient payload (model declined to judge) | 0.004 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
The models disagree on parts of this classification; every voice is preserved in the section at the end of the page.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".