Contrasting performance of panel and time-series data models for subnational crop forecasting in Sub-Saharan Africa
Bibliographic record
Abstract
• Panel and time-series data models are compared for crop production and yield forecasting. • Panel data model provides yield predictions comparable to time-series data model. • Time-invariant features capture spatial variability, enhancing climatic modeling. • Longer training boosts panel model's adaptability to production outliers. • We advocate for region-specific methods to consider spatiotemporal nuances. We comprehensively examine methodologies tailored for subnational crop yield and production forecasting by integrating Earth Observation (EO) datasets and advanced machine learning approaches. We scrutinized diverse input data types, cross-validation methods, and training durations, focusing on maize production and yield predictions in Burkina Faso and Somalia. Central to our analysis is the comparative assessment of using time-invariant features within a panel data (PD) model versus a time-series data (TD) model. The TD model performed well in predicting both production and yield, while the PD model offered comparable yield predictions. Time-invariant features such as livelihood zones, soil properties, and cropland extents enriched the spatial understanding of crop data, enhancing the R-squared by 0.09 (0.21) for production and 0.11 (0.03) for yield, with corresponding reductions in the Mean Absolute Percentage Error by 90 % (238 %) for production and 5 % (4 %) for yield in Burkina Faso (Somalia). While Burkina Faso's consistent crop data allowed for effective modeling with brief training, Somalia benefited from the adaptability of the PD model to crop statistics outliers, particularly with extended training in high-producing regions. The PD approach showed promise in addressing data gaps, although predicting crop productions for unobserved districts remained a challenge. Our findings highlight the harmonious integration of EO data and machine learning in the field of agricultural forecasting and emphasize the importance of region-specific methodologies, especially in the rapidly changing landscape of EO data convergence.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.006 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".