Using monitoring and evaluation to build equity and resilience: lessons from practice
Bibliographic record
Abstract
The field of monitoring and evaluation (M&E) is intimately connected with issues of power. Power is exercised in choices regarding what is monitored and evaluated; by, for, and with whom this is done; how data are collected; which criteria are used to indicate success; with whom results are shared and for what purpose; and who learns what in the process. M&E findings play a crucial role in determining whether funding and support for initiatives and organizations are continued or stopped. Therefore, the way in which M&E is practiced can profoundly influence whether it promotes equity and resilience or, conversely, dominance, exclusion, and dependency. This paper presents four insights into how M&E practice can contribute to building equity and resilience. These insights are drawn from the authors’ reflections on their experiences as practitioners, facilitated through participation in a Southern African Resilience Academy M&E working group. The working group provided an opportunity to shift practice into knowledge, contrasting with the more commonly used concept of shifting knowledge into practice. Six case studies were used to reflect on successful and unsuccessful aspects within the often messy, contested, and resource-limited contexts of organizations and projects. The paper identifies possible systemic leverage points for building transformative equity and resilience through M&E.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.101 | 0.088 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.004 | 0.003 |
| Science and technology studies | 0.006 | 0.035 |
| Scholarly communication | 0.011 | 0.015 |
| Open science | 0.003 | 0.016 |
| Research integrity | 0.004 | 0.006 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".