Impact Evaluation of Water Infrastructure Investments: Methods, Challenges and Demonstration From a Large‐Scale Urban Improvement in Jordan
Bibliographic record
Abstract
Abstract Impact evaluation (IE) of large infrastructure presents numerous challenges, and investments in urban piped water and sanitation are no exception. Here we present methods for more systematic assessment of the implications of such interventions, discussing tradeoffs between validity, relevance and practicality that arise from alternative approaches. Then, to more clearly illustrate the many issues that typically arise in such IEs, we draw on an example application in Zarqa, Jordan, where the Millennium Challenge Corporation invested about US$275 million to upgrade and extend piped water and sewer networks, as well as increase the capacity of the country's largest wastewater treatment plant. The theory of change for the intervention took a systems view of impacts: the project aimed to improve water supply to urban areas while maintaining flows to irrigators through enhanced wastewater reuse. The case adds valuable evidence on the impacts of large infrastructure investments and illustrates well the challenges of capturing spillovers, mitigating study contamination, maintaining statistical power, and determining overall welfare effects, in situations involving diverse market and nonmarket impacts. These limitations notwithstanding, the application highlights the high value of conducting IEs, and why applied researchers should not give up on pragmatic and interdisciplinary collaborations to evaluation in the face of complex interventions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".