The global stock of research evidence relevant to health systems policymaking
Bibliographic record
Abstract
BACKGROUND: Policymakers and stakeholders need immediate access to many types of research evidence to make informed decisions about the full range of questions that may arise regarding health systems. METHODS: We examined all types of research evidence about governance, financial and delivery arrangements, and implementation strategies within health systems contained in Health Systems Evidence (HSE) (http://www.healthsystemsevidence.org). The research evidence types include evidence briefs for policy, overviews of systematic reviews, systematic reviews of effects, systematic reviews addressing other questions, systematic reviews in progress, systematic reviews being planned, economic evaluations, and health reform and health system descriptions. Specifically, we describe their distribution across health system topics and domains, trends in their production over time, availability of supplemental content in various languages, and the extent to which they focus on low- and middle-income countries (LMICs), as well as (for systematic reviews) their methodological quality and the availability of user-friendly summaries. RESULTS: As of July 2013, HSE contained 2,629 systematic reviews of effects (of which 501 are Cochrane reviews), 614 systematic reviews addressing other questions, 283 systematic reviews in progress, 186 systematic reviews being planned, 140 review-derived products (evidence briefs and overviews of systematic reviews), 1,669 economic evaluations, 1,092 health reform descriptions, and 209 health system descriptions. Most systematic reviews address topics related to delivery arrangements (n = 2,663) or implementation strategies (n = 1,653) with far fewer addressing financial (n = 241) or governance arrangements (n = 231). In addition, 2,928 systematic reviews have been quality appraised with moderate AMSTAR ratings found for reviews addressing governance (5.6/11), financial (5.9/11), and delivery (6.3/11) arrangements and implementation strategies (6.5/11); 1,075 systematic reviews have no independently produced user-friendly summary and only 737 systematic reviews have an LMIC focus. Literature searches for half of the systematic reviews (n = 1,584, 49%) were conducted within the last five years. CONCLUSIONS: Greater effort needs to focus on assessing whether the current distribution of systematic reviews corresponds to policymakers' and stakeholders' priorities, updating systematic reviews, increasing the quality of systematic reviews, and focusing on LMICs.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.132 | 0.285 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.005 | 0.004 |
| Bibliometrics | 0.045 | 0.064 |
| Science and technology studies | 0.002 | 0.012 |
| Scholarly communication | 0.030 | 0.025 |
| Open science | 0.004 | 0.013 |
| Research integrity | 0.011 | 0.010 |
| Insufficient payload (model declined to judge) | 0.030 | 0.007 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".