Evidence for Resilient Agriculture Dataset
Bibliographic record
Abstract
The Evidence for Resilient Agriculture (ERA) dataset now synthesizes evidence from 2,916 agricultural studies conducted across Africa, providing a comprehensive foundation for evaluating the performance of agronomic technologies and management strategies in diverse contexts. ERA v1.0.1 contains 112,859 observations from 2,011 agricultural studies published between 1934 and 2018. These studies examine the efficacy of 363 practice combinations across 87 environmental, social, and agricultural-economic outcome indicators. Observations are geolocated and can be linked to open-source environmental, economic, and social datasets, enabling analysis of how local context shapes the performance of agricultural practices. ERA v1 provides a foundational evidence base for the design of policies, programs, and investments supporting African agricultural development. As part of the 2024–25 update (ERA v2), we expanded the dataset to include studies published between 2018 and 2024. This additional search identified approximately 900 new eligible studies, bringing the total number of studies represented in ERA to 2,916. The expansion substantially strengthens the evidence base for understanding climate-resilient agronomy across Africa, particularly in areas such as soil fertility management, climate adaptation, intercropping, and sustainable intensification. In addition to extending the temporal coverage, ERA v2 modernized and enhanced its methodology through the integration of AI-powered tools. Literature screening was supported by OpenAlex, an open research graph enabling scalable, automated discovery and filtering of scientific publications (see vignette: https://eragriculture.github.io/AI-Powered-Meta-Analysis-Automation/docs/OA-vignette.html). Data extraction workflows were augmented using OpenAI APIs, which support semi-automated extraction of numerical results and metadata from tables, text, and figures (https://eragriculture.github.io/AI-Powered-Meta-Analysis-Automation/docs/Use_of_AI_for_Extraction.html). These innovations substantially increased throughput, consistency, and reproducibility in evidence synthesis, reducing human extraction time while improving dataset structure and reliability. The ERA dataset includes bibliographic metadata, geographic coordinates, environmental context, experimental design variables, treatment comparisons, and outcome indicators. Each row corresponds to a unique combination of article, site, treatment contrast, commodity, outcome, and time period. Supporting documentation provides definitions, hierarchies, and data structures for all coded fields: -ERA_Compiled.csv – compiled ERA dataset (wide format) -ERA_Compiled_Fields.csv – descriptions of dataset fields -ERA_Bibliography.csv – bibliographic metadata -ERA_Search_Terms.csv – search terms for literature discovery -Practice_Codes.csv – hierarchical definitions of agronomic practices -Outcome_Codes.csv – outcome definitions and hierarchies -EU_Codes.csv – enterprise unit definitions To support users, a fully updated ERA User Guide has been published: https://eragriculture.github.io/ERA_Agronomy/ERA-User-Guide.html https://eragriculture.github.io/ERL/Guide-to-Livestock-Data-Analysis-in-the-ERA-Dataset--STATIC.html Additional vignettes from the ERAg and ERAgON R packages illustrate workflows for data exploration, analysis, and reproducibility: -ERA-Introduction.pdf -ERAdev – a collection of scripts illustrating how the ~2,900 studies in ERA were systematically compiled, standardized, and harmonized (AI-assisted workflows included) -ERA-Explore-and-Analyze.pdf -ERA-Search-Protocols.pdf -ERA-Yield-Stability.pdf ERA v1 have received support from the CGIAR Excellence in Agronomy Initiative, the Livestock and Climate Initiative, the CGIAR Research Program on Climate Change, Agriculture, and Food Security (CCAFS), and partner agencies including FAO, the EU, IFAD, USDA-FAS, and CIFOR’s Evidence-Based Forestry program. The ERA v2 update was additionally funded by the CGIAR Sustainable Farming Program (SFP).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.014 | 0.118 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.003 | 0.004 |
| Bibliometrics | 0.018 | 0.023 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.003 | 0.005 |
| Research integrity | 0.003 | 0.003 |
| Insufficient payload (model declined to judge) | 0.070 | 0.008 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".