Enhancing Winter Wheat Yield Estimation Using Machine Learning and Fusion of Radar and Optical Satellite Imagery
Bibliographic record
Abstract
Accurate crop yield Mapping is paramount in agricultural monitoring and food security. In this study, we present a comprehensive investigation into estimating winter wheat yield in the Qazvin plane of Iran, leveraging the synergy between machine learning algorithms and the fusion of remote sensing data from radar and optical satellite sensors. The research is based on the availability of high-quality in situ yield data gathered by the Ministry of Agriculture in collaboration with the Food and Agriculture Organization (FAO), collected during the 2019–2020 crop year. The study area encompasses the Qazvin plane, an agriculturally significant region renowned for winter wheat production in Iran. In-situ data from various agricultural fields and seed types as reference measurements enabled us to conduct rigorous validation of the performance of machine learning algorithms and the effectiveness of the fused remote sensing data. The primary objective of this study is to assess and compare the performance of seven prominent machine learning algorithms for accurate estimation of the annual winter wheat yields. Furthermore, we investigate the individual and synergistic capabilities of radar and optical satellite sensors in estimating winter wheat yield. Through rigorous analysis of the pixel-level confusion matrices, we identify the most effective model for yield estimation, evaluating the complementarity and information redundancy between the two types of remote sensing data. In this study, we conducted an extensive comparison of various machine learning algorithms for winter wheat crop yield estimation in the Qazvin plane of Iran. Among the four best-performing algorithms examined, namely polynomial regression (RMSE = 0.5657 t/ha−1), random forest (RMSE = 0.1632 t/ha−1), XGBoost (RMSE = 0.3153 t/ha−1), and the proposed Multi-Layer Perceptron (MLP) (RMSE = 0.1324 t/ha−1), the MLP demonstrated superior performance. The MLP’s yield estimation exceeded the total yearly agricultural statistics of Qazvin by 0.19 percent. However, this discrepancy can be attributed to various factors, including errors in wheat and barley field mapping, miscalculation in cumulative statistics, and the inherent limitations of yield estimation algorithms in capturing the dynamic nature of agricultural systems. The findings of this research provide valuable insights into the potential of machine learning algorithms and remote sensing data fusion for accurate crop yield estimation, paving the way for enhanced agricultural monitoring and decision-making processes in the region.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".