Tackling the tensions in evaluating capacity strengthening for health research in low- and middle-income countries
Bibliographic record
Abstract
Strengthening research capacity in low- and middle-income countries is one of the most effective ways of advancing their health and development but the complexity and heterogeneity of health research capacity strengthening (RCS) initiatives means it is difficult to evaluate their effectiveness. Our study aimed to enhance understanding about these difficulties and to make recommendations about how to make health RCS evaluations more effective. Through discussions and surveys of health RCS funders, including the ESSENCE on Health Research initiative, we identified themes that were important to health RCS funders and used these to guide a systematic analysis of their evaluation reports. Eighteen reports, produced between 2000 and 2013, representing 12 evaluations, were purposefully selected from 54 reports provided by the funders to provide maximum variety. Text from the reports was extracted independently by two authors against a pre-designed framework. Information about the health RCS approaches, tensions and suggested solutions was re-constructed into a narrative. Throughout the process contacts in the health RCS funder agencies were involved in helping us to validate and interpret our results. The focus of the health RCS evaluations ranged from individuals and institutions to national, regional and global levels. Our analysis identified tensions around how much stakeholders should participate in an evaluation, the appropriate balance between measuring and learning and between a focus on short-term processes vs longer-term impact and sustainability. Suggested solutions to these tensions included early and ongoing stakeholder engagement in planning and evaluating health RCS, modelling of impact pathways and rapid assimilation of lessons learned for continuous improvement of decision making and programming. The use of developmental approaches could improve health RCS evaluations by addressing common tensions and promoting sustainability. Sharing learning about how to do robust and useful health RCS evaluations should happen alongside, not after, health RCS efforts.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".