Dynamic evaluation of agricultural research for development supports innovation and responsible scaling through high-level inclusion
Bibliographic record
Abstract
Innovators want to scale the impact of innovations responsibly, but the meaning of “responsible” is elusive. It entails the high-level inclusion of stakeholders, but there is no agreed-upon standard that defines a “good-enough” level. Our objectives are to (1) provoke a conversation about what it means to scale the impact of agricultural innovations responsibly and (2) suggest dynamic evaluation as one way to promote responsible scaling because it facilitates the leadership of the people affected. Over 200 projects funded by IDRC to scale the impact of research for development in the Global South were reviewed. Research products were iteratively developed with Southern innovators and Northern funders who offered structured feedback. The dynamic scaling systems model is one research product. It can guide the people affected as they lead a process of “conjecturing” about scaling effects in complex settings. The resulting conjectures inform the dynamic evaluation of scaling, as well as planning and management. We illustrate the application of the model with a hypothetical and real example. Scaling is an integral part of agricultural innovation, and dynamism is an emerging concept that informs the evaluation, planning, and management of scaling. The dynamic scaling systems model supports the high-level inclusion of the people affected in ways that respect local knowledge and the risks associated with complex settings. It helps innovators scale more responsibly, even though the precise meaning of “responsible” remains elusive. • Innovators want to scale the impacts of agricultural innovations responsibly, but the meaning of “responsible” is elusive. • Innovators cannot fully achieve ideal inclusion—everyone affected participates at the highest level with complete control. • Dynamic evaluation helps innovators get closer to the ideal because, among other things, it promotes Conjecturing . • Conjecturing is a collaborative process of anticipating scaling effects that is led by the people affected. • The dynamic scaling systems model guides conjecturing by framing questions about potential scaling effects in complex systems.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.187 | 0.208 |
| Meta-epidemiology (narrow) | 0.002 | 0.001 |
| Meta-epidemiology (broad) | 0.002 | 0.002 |
| Bibliometrics | 0.006 | 0.004 |
| Science and technology studies | 0.009 | 0.038 |
| Scholarly communication | 0.034 | 0.029 |
| Open science | 0.004 | 0.029 |
| Research integrity | 0.007 | 0.007 |
| Insufficient payload (model declined to judge) | 0.009 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".