We need a NICE for global development spending
Bibliographic record
Abstract
<ns4:p> With aid budgets shrinking in richer countries and more money for healthcare becoming available from domestic sources in poorer ones, the rhetoric of value for money or improved efficiency of aid spending is increasing. Taking healthcare as one example, we discuss the need for and potential benefits of (and obstacles to) the establishment of a national institute for aid effectiveness. In the case of the UK, such an institute would help improve development spending decisions made by DFID, the country’s aid agency, as well as by the various multilaterals, such as the Global Fund, through which British aid monies is channelled. It could and should also help countries becoming increasingly independent from aid build their own capacity to make sure their own resources go further in terms of health outcomes and more equitable distribution. Such an undertaking will not be easy given deep suspicion amongst development experts towards economists and arguments for improving efficiency. We argue that it is exactly <ns4:italic>because</ns4:italic> needs matter that those who make spending decisions must consider the needs not being met when a priority requires that finite resources are diverted elsewhere. These chosen unmet needs are the true costs; they are lost health. They <ns4:italic>must</ns4:italic> be considered, and should be minimised and must therefore be measured. Such exposition of the trade-offs of competing investment options can help inform an array of old and newer development tools, from strategic purchasing and pricing negotiations for healthcare products to performance based contracts and innovative financing tools for programmatic interventions. </ns4:p>
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".