Estimating crop type and yield of small holder fields in Burkina Faso using multi-day Sentinel-2
Bibliographic record
Abstract
Remote Sensing affords the opportunity to monitor and evaluate data scarce regions where field collection efforts are costly. A particular challenge is monitoring and evaluation in regions with smallholder agricultural systems (∼1 ha) that are often subsistence focused, vulnerable to food insecurity and data scarce. Using multi-day moderate resolution Sentinel-2 and Random Forest models, this study shows that crop type and rice yields in Burkina Faso can be predicted with greater than ∼80% accuracy in the rainy season. Model optimization using varying spectral and vegetation index inputs can increase crop type and yield prediction accuracy in the dry season where denser cultivation is a challenge for the 10–20 m resolution of Sentinel-2. However, there is a trade-off between opting for very high-resolution imagery (<2 m) or the number of bands offered by Sentinel-2 as the bands that occupy and vegetation indices that utilize the red through NIR ranges were most important across all models. In addition, model type, linear Regression or nonlinear Random Forest, matters little when estimating yield in these landscapes, unless Harmonic regression is utilized for the linear model. This study also showed that a model trained with high quality 2019 dry season crop cut data can predict the subsequent dry season's interannual crop type with overall accuracy as high as 60%, comparable to crop type models trained with 2020 survey data and used to estimate crop type in the concurrent season, as the survey collection. This indicates some utility in leveraging the calibrated Random Forest models to make skillful predictions of interannual crop type and ultimately food availability of nearby communities for years with no training data. Given increasing global food prices and restricted commodity trade, understanding local agricultural productivity using affordable and timely remote sensing-based methods is essential for ensuring appropriate humanitarian interventions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".