Fostering Responsible Innovation in Health: An EvidenceInformed Assessment Tool for Innovation Stakeholders
Bibliographic record
Abstract
BACKGROUND: Responsible innovation in health (RIH) emphasizes the importance of developing technologies that are responsive to system-level challenges and support equitable and sustainable healthcare. To help decision-makers identify whether an innovation fulfills RIH requirements, we developed and validated an evidence-informed assessment tool comprised of 4 inclusion and exclusion criteria, 9 assessment attributes and a scoring system. METHODS: We conducted an inter-rater reliability assessment to establish the extent to which 2 raters agree when applying the RIH Tool to a diversified sample of health innovations (n=25). Following the Tool's 3-step process, sources of information were collected and cross-checked to ensure their clarity and relevance. Ratings were reported independently in a spreadsheet to generate the study's database. To measure inter-rater reliability, we used: a non-adjusted index (percent agreement), a chance-adjusted index (Gwet's AC) and the Pearson's correlation coefficient. Results of the Tool's application to the whole sample of innovations are summarized through descriptive statistics. RESULTS: Our findings show complete agreement for the screening criteria, "almost perfect" agreement for 7 assessment attributes, "substantial" agreement for 2 attributes and "almost perfect" agreement for the RIH overall score. A large portion of the sample obtained high scores for 6 attributes (health relevance, health inequalities, responsiveness, level and intensity of care and frugality) and low scores for 3 attributes (ethical, legal, and social issues [ELSIs], inclusiveness and eco-responsibility). At the rating step, 88% of the innovations had a sufficient number of attributes documented (≥ 7/9), but the assessment was based on sources of moderate to high quality (mean score ≥ 2 points) for 36% of the sample. While "Almost all RIH features" were present for 24% of the innovations (RIH mean score between 4.1-5.0 points), "Many RIH features" were present for 52% of the sample (3.1-4.0 points) and "Few RIH features" were present for 24% of the innovations (2.1-3.0 points). CONCLUSION: By confirming key aspects of the RIH Tool's reliability and applicability, our study brings its development to completion. It can be jointly put into action by innovation stakeholders who want to foster innovations with greater social, economic and environmental value.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.214 | 0.364 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.006 |
| Bibliometrics | 0.028 | 0.018 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.011 | 0.012 |
| Open science | 0.005 | 0.016 |
| Research integrity | 0.004 | 0.003 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".