Toward improving global rice yield reference dataset compilation through machine learning: Insights from training data selection and random forest analysis
Bibliographic record
Abstract
Machine learning (ML) techniques have been increasingly used to estimate crop yields at scales ranging from on-site to global. Since ML techniques are data-driven approaches, it is empirically known that the performance of a specific ML algorithm depends on the manner in which the training dataset is compiled. However, few studies have quantitatively evaluated the performance. In this study, global rice yields were estimated through a random forest (RF) methodology. Performance dependency of RF on training data was examined by a comparison of estimated yields using different training datasets covering different yield ranges and geographical extents. First, 14 explanatory variables collected from different sources (satellite vegetation, meteorology, and geographical location data) were used for building RF regressors. The crop calendar was determined from a combination of satellite vegetation and crop model simulation. Next, RF regressors were trained to give census-based rice yields (used as reference yields) from training datasets of the 14 explanatory variables. By applying the RF regressors to validation datasets, misfits between estimated and the reference yields were evaluated. RF reproduced rice yields, but the accuracy depended on the training data. Yields beyond the yield range of the training data could not be reproduced by RF. This indicates that the yield range of the training data determined the possible range of estimated yield. Among the 14 variables, geographical coordinates (longitude and latitude) ranked the highest importance, i.e., played a crucial role in estimating yields. The RF regressors built from the 14 variables outperformed those built only from the geographical coordinates in accuracy but with limited advantage. We concluded that (1) choosing training data to cover all possible yield ranges of the target rice-cropping areas was crucial for accurate yield estimation using RF and (2) incorporating satellite and simulation data was advantageous for building high-performance RF regressors.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".