Evaluating research co-production: protocol for the Research Quality Plus for Co-Production (RQ+ 4 Co-Pro) framework
Bibliographic record
Abstract
BACKGROUND: Research co-production is an umbrella term used to describe research users and researchers working together to generate knowledge. Research co-production is used to create knowledge that is relevant to current challenges and to increase uptake of that knowledge into practice, programs, products, and/or policy. Yet, rigorous theories and methods to assess the quality of co-production are limited. Here we describe a framework for assessing the quality of research co-production-Research Quality Plus for Co-Production (RQ+ 4 Co-Pro)-and outline our field test of this approach. METHODS: Using a co-production approach, we aim to field test the relevance and utility of the RQ+ 4 Co-Pro framework. To do so, we will recruit participants who have led research co-production projects from the international Integrated Knowledge Translation Research Network. We aim to sample 16 to 20 co-production project leads, assign these participants to dyadic groups (8 to 10 dyads), train each participant in the RQ+ 4 Co-Pro framework using deliberative workshops and oversee a simulation assessment exercise using RQ+ 4 Co-Pro within dyadic groups. To study this experience, we use a qualitative design to collect participant demographic information and project demographic information and will use in-depth semi-structured interviews to collect data related to the experience each participant has using the RQ+ 4 Co-Pro framework. DISCUSSION: This study will yield knowledge about a new way to assess research co-production. Specifically, it will address the relevance and utility of using RQ+ 4 Co-Pro, a framework that includes context as an inseparable component of research, identifies dimensions of quality matched to the aims of co-production, and applies a systematic and transferable evaluative method for reaching conclusions. This is a needed area of innovation for research co-production to reach its full potential. The findings may benefit co-producers interested in understanding the quality of their work, but also other stewards of research co-production. Accordingly, we undertake this study as a co-production team representing multiple perspectives from across the research enterprise, such as funders, journal editors, university administrators, and government and health organization leaders.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Direct model labels (unvalidated)
Per-model category and study-design labels from the labeling rounds. They are machine output, unvalidated, and the disagreement between models ships as data. No study design here is MEDLINE-validated yet.
| Model arm | Categories | Study design | Confidence |
|---|---|---|---|
| gemma | Metaresearch Domain: Evaluation · Genre: Protocol About the Canadian research system: no · About a Canadian topic: no | Not applicable | low |
| gpt | Metaresearch Domain: Evaluation · Genre: Protocol About the Canadian research system: no · About a Canadian topic: no | Not applicable | high |
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.269 | 0.385 |
| Meta-epidemiology (narrow) | 0.005 | 0.005 |
| Meta-epidemiology (broad) | 0.007 | 0.008 |
| Bibliometrics | 0.007 | 0.011 |
| Science and technology studies | 0.009 | 0.008 |
| Scholarly communication | 0.010 | 0.009 |
| Open science | 0.006 | 0.011 |
| Research integrity | 0.010 | 0.015 |
| Insufficient payload (model declined to judge) | 0.098 | 0.036 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedLabeled directly by 2 models reading the full record.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".