Bibliographic record
Abstract
Time to act Research misconduct can harm patients, distort the evidence base, misdirect research effort, waste funds, and damage public trust in science. Countries all over the developed world are now recognising the need to set up systems to deter, detect, and investigate research misconduct. Why does the United Kingdom have no plans to do the same? As Aniket Tavare outlines in the linked feature (doi:10.1136/bmj.d8212),1 high profile cases of misconduct have led the United States, Canada, Sweden, Norway, and Poland, among others, to create formal mechanisms for overseeing research integrity. In most countries responsibility lies with the institutions, but oversight varies greatly, and it is unclear which systems are most effective and efficient. None is perfect—the remit of the US Office of Research Integrity is limited to publicly funded health research; Australia’s recently established Research Integrity Committee is already being criticised for lacking teeth. But each system shows that the problem has been acknowledged, that institutions accept primary responsibility, and that governments and funders are seriously committed to tackling misconduct openly and with a range of statutory powers. In contrast, the UK has no official national body. The UK Research Integrity Office was established in 2006 and has done some useful things. But its function has always been advisory, and now that the major funders represented by Research Councils UK (RCUK) have decided not to continue the funding, it relies on voluntary funding from institutions. The Research Integrity Futures Working Group, set up by RCUK and Universities UK (UUK) and other bodies, has also apparently come to nothing. The working group’s report commissioned in 2009 called for an independent advisory body, …
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | MetaresearchResearch integrity Domain: Methods · Genre: Commentary About the Canadian research system: no · About a Canadian topic: no | Not applicable | high |
| gpt | Research integrity Domain: not available · Genre: Commentary About the Canadian research system: no · About a Canadian topic: no | Not applicable | high |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.055 | 0.331 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.006 | 0.006 |
| Science and technology studies | 0.009 | 0.015 |
| Scholarly communication | 0.020 | 0.013 |
| Open science | 0.004 | 0.017 |
| Research integrity | 0.024 | 0.016 |
| Insufficient payload (model declined to judge) | 0.049 | 0.014 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
The models disagree on parts of this classification; every voice is preserved in the section at the end of the page.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".