MétaCan
Menu
Back to cohort
Record W4406279610 · doi:10.1016/j.jag.2025.104367

The impact of spatiotemporal variability of environmental conditions on wheat yield forecasting using remote sensing data and machine learning

2025· article· en· W4406279610 on OpenAlexaff
Keltoum Khechba, Mariana Belgiu, Ahmed Laamrani, Alfred Stein, Abdelhakim Amazirh, Abdelghani Chehbouni

Bibliographic record

VenueInternational Journal of Applied Earth Observation and Geoinformation · 2025
Typearticle
Languageen
FieldEnvironmental Science
TopicRemote Sensing in Agriculture
Canadian institutionsUniversity of Guelph
FundersUniversité Mohammed VI Polytechnique
KeywordsYield (engineering)GeographyRemote sensingEnvironmental science

Abstract

fetched live from OpenAlex

• Zone-based ML for field-level yield forecasting using environmental and remote sensing variables. • Spatial and temporal variability of the input variables affect ML accuracy. • Field management and weather shifts cause monthly fluctuations in ML accuracy. Climate change poses significant challenges to food security, especially in semi-arid agriculture areas. Effective monitoring of crop yield is important for establishing food emergency responses and developing long-term sustainable strategies. In Morocco, where cereals are the predominant crops, yield forecasting is important for addressing the yield gap as it enables farmers to take preventive actions before the harvesting period. This study aims to assess the impact of spatial and temporal heterogeneity of environmental conditions on wheat yield forecasting using machine learning models. It compares the 2019–2020 and 2020–2021 agricultural seasons using three sets of variables: (1) spectral indices; (2) weather data; and (3) a combination of both spectral indices and weather data. Weather data, including cumulative monthly precipitation from ERA5 data and average monthly temperature from PERSIANN data, were extracted for the wheat growing season (November to June). Spectral indices including the Normalized Difference Vegetation Index, Moisture Stress Index, and Terrestrial Chlorophyll Index were calculated from Sentinel-2 imagery for the same period and processed using Google Earth Engine. The study area was divided into homogeneous zones based on an existing landform classification, and XGBoost and Random Forest (RF) models were used for yield forecasting in each zone separately. The two models performed equally well across both the zones and the whole study area (SA) when using weather data as the input variable. For instance, across SA, they achieved average R 2 values of 0.60 and 0.81 for all months during the 2019–2020 and 2020–2021 agricultural seasons, respectively. However, when using spectral indices or combining these indices with weather data, RF consistently outperformed XGBoost. For example, in SA during the 2019–2020 season, RF achieved an average R 2 of 0.48 across the growing season, compared to XGBoost’s R 2 of 0.43. Similarly, in the 2020–2021 season, RF achieved an R 2 of 0.35 and an RMSE of 1083 kg ha -1 , while XGBoost performed slightly lower, with an R 2 of 0.29 and an RMSE of 1137 kg ha -1 . Comparing the prediction accuracy between the seasons for each set of variables, the RF model performs better when using spectral indices during the relatively dry 2019–2020 season as compared to the wet 2020–2021 season. Incorporating weather data, the model improved its performance for the 2020–2021 season. April showed the highest prediction performance overall, with R 2 values of 0.6 for SA using weather data alone in the 2019–2020 season, and 0.8 for SA using a combination of weather data and spectral indices in the 2020–2021 season. The 2019–2020 season showed strong fluctuations in accuracy throughout the growing season, whereas the 2020–2021 season had a consistent improvement in accuracy over time. These variations in accuracy are due to differing environmental conditions that should be taken into account for making better and more reliable yield predictions.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.002
metaresearch head score (Gemma)0.005
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.020
Threshold uncertainty score0.040

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0020.005
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.001
Bibliometrics0.0010.001
Science and technology studies0.0000.000
Scholarly communication0.0010.001
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.033
GPT teacher head0.261
Teacher spread0.228 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations9
Published2025
Admission routes1
Has abstractyes

Explore more

Same venueInternational Journal of Applied Earth Observation and GeoinformationSame topicRemote Sensing in AgricultureFrench-language works237,207