Holy grail or convenient excuse? Stakeholder perspectives on the role of health system strengthening evaluation in global health resource allocation
Bibliographic record
Abstract
BACKGROUND: The role of evaluation evidence in guiding health systems strengthening (HSS) investments at the global-level remains contested. A lack of rigorous impact evaluations is viewed by some as an obstacle to scaling resources. However, others suggest that power dynamics and knowledge hierarchies continue to shape perceptions of rigor and acceptability in HSS evaluations. This debate has had major implications on HSS resource allocation in global-level funding decisions. Yet, few studies have examined the relationship between HSS evaluation evidence and prioritization of HSS. In this paper, we explore the perspectives of key global health stakeholders, specifically around the nature of evidence sought regarding HSS and its potential impact on prioritization, the challenges in securing such evidence, and the drivers of intra- and inter-organizational divergences. We conducted a stakeholder analysis, drawing on 25 interviews with senior representatives of major global health organizations, and utilized inductive approaches to data analysis to develop themes. RESULTS: Our analysis suggests an intractable challenge at the heart of the relationship between HSS evaluations and prioritization. A lack of evidence was used as a reason for limited investments by some respondents, citing their belief that HSS was an unproven and potentially risky investment which is driven by the philosophy of HSS advocates rather than evidence. The same respondents also noted that the 'holy grail' of evaluation evidence that they sought would be rigorous studies that assess the impact of investments on health outcomes and financial accountability, and believed that methodological innovations to deliver this have not occurred. Conversely, others held HSS as a cross-cutting principle across global health investment decisions, and felt that the type of evidence sought by some funders is unachievable and not necessary - an 'elusive quest' - given methodological challenges in establishing causality and attribution. In their view, evidence would not change perspectives in favor of HSS investments, and evidence gaps were used as a 'convenient excuse'. Respondents raised additional concerns regarding the design, dissemination and translation of HSS evaluation evidence. CONCLUSIONS: Ongoing debates about the need for stronger evidence on HSS are often conducted at cross-purposes. Acknowledging and navigating these differing perspectives on HSS evaluation may help break the gridlock and find a more productive way forward.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".