Academic contributions to the development of evidence and policy systems: an EPPI Centre collective autoethnography
Bibliographic record
Abstract
BACKGROUND: Evidence for policy systems emerging around the world combine the fields of research synthesis, evidence-informed policy and public engagement with research. We conducted this retrospective collective autoethnography to understand the role of academics in developing such systems. METHODS: We constructed a timeline of EPPI Centre work and associated events since 1990. We employed: Transition Theory to reveal emerging and influential innovations; and Transformative Social Innovation theory to track their increasing depth, reach and embeddedness in research and policy organisations. FINDINGS: The EPPI Centre, alongside other small research units, collaborated with national and international organisations at the research-policy interface to incubate, spread and embed new ways of working with evidence and policy. Sustainable change arising from research-policy interactions was less about uptake and embedding of innovations, but more about co-developing and tailoring innovations with organisations to suit their missions and structures for creating new knowledge or using knowledge for decisions. Both spreading and embedding innovation relied on mutual learning that both accommodated and challenged established assumptions and values of collaborating organisations as they adapted to closer ways of working. The incubation, spread and embedding of innovations have been iterative, with new ways of working inspiring further innovation as they spread and embedded. Institutionalising evidence for policy required change in both institutions generating evidence and institutions developing policy. CONCLUSIONS: Key mechanisms for academic contributions to advancing evidence for policy were: contract research focusing attention at the research-policy interface; a willingness to work in unfamiliar fields; inclusive ways of working to move from conflict to consensus; and incentives and opportunities for reflection and consolidating learning.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | Metaresearch Domain: Methods · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Qualitative | low |
| gpt | Metaresearch Domain: Methods · Genre: Empirical About the Canadian research system: no · About a Canadian topic: no | Qualitative | low |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.070 | 0.086 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.004 |
| Science and technology studies | 0.019 | 0.036 |
| Scholarly communication | 0.012 | 0.012 |
| Open science | 0.003 | 0.024 |
| Research integrity | 0.004 | 0.010 |
| Insufficient payload (model declined to judge) | 0.005 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".