Alfalfa stem count estimation using remote sensing imagery and machine learning on Google Earth Engine
Bibliographic record
Abstract
Alfalfa ( Medicago sativa L. ), a perennial legume forage crop, is valued for its high yield and quality. However, its survival during winter can be affected by several factors, and its mortality significantly impacts alfalfa production, necessitating timely and spatially detailed monitoring. This study aims to propose a framework for estimating alfalfa stem density using satellite imagery and machine learning (ML) algorithms, which can lead to winter mortality detection early in the spring and provide a better understanding of potential total dry matter. Three ML models—support vector machine (SVM), random forest (RF), and extreme gradient boosting (XGB)—were applied to Harmonized Landsat Sentinel (Landsat only, which is HLSL30) and Sentinel-2 datasets, accessed via the Google Earth Engine (GEE) Python API. Two scenarios were evaluated: 1) single-date data, capturing satellite images within a 3-day time window to the date of field sample measurement, and 2) time-series data, in which three satellite images were collected for the measurements during the first growing cycle. Both classification and regression models were used in both scenarios to estimate and classify alfalfa stem density. ML Classification models categorized stem density into four groups (bare, low-density, medium-density, and high-density), achieving an accuracy of up to 85 % using Sentinel-2 data and 84 % using HLSL30 data. The results also indicated that alfalfa stem density can be estimated with an error of ∼ ±6-9 stems/foot 2 (1 foot = 30.48 cm) using ML regression models. RF outperformed XGB and SVM in classification and regression tasks, showing superior accuracy in classifying density and lower root mean square error (RMSE) in estimating stem density. Our proposed framework model can offer valuable information to growers and decision-makers, enabling them to make timely and informed decisions.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".