Integrating the Strength of Multi-Date Sentinel-1 and -2 Datasets for Detecting Mango (Mangifera indica L.) Orchards in a Semi-Arid Environment in Zimbabwe
Bibliographic record
Abstract
Generating tree-specific crop maps within heterogeneous landscapes requires imagery of fine spatial and temporal resolutions to discriminate among the rapid transitions in tree phenological and spectral features. The availability of freely accessible satellite data of relatively high spatial and temporal resolutions offers an unprecedented opportunity for wide-area land use and land cover (LULC) mapping, including tree crop (e.g., mango; Mangifera indica L.) detection. We evaluated the utility of combining Sentinel-1 (S1) and Sentinel-2 (S2) derived variables (n = 81) for mapping mango orchard occurrence in Zimbabwe using machine learning classifiers, i.e., support vector machine and random forest. Field data were collected on mango orchards and other LULC classes. Fewer variables were selected from ‘All’ combined S1 and S2 variables using three commonly utilized variable selection methods, i.e., relief filter, guided regularized random forest, and variance inflation factor. Several classification experiments (n = 8) were conducted using 60% of field datasets and combinations of ‘All’ and fewer selected variables and were compared using the remaining 40% of the field dataset and the area underclass approach. The results showed that a combination of random forest and relief filter selected variables outperformed (F1 score > 70%) all other variable combination experiments. Notwithstanding, the differences among the mapping results were not significant (p ≤ 0.05). Specifically, the mapping accuracy of the mango orchards was more than 80% for each of the eight classification experiments. Results revealed that mango orchards occupied approximately 18% of the spatial extent of the study area. The S1 variables were constantly selected compared with the S2-derived variables across the three variable selection approaches used in this study. It is concluded that the use of multi-modal satellite imagery and robust machine learning classifiers can accurately detect mango orchards and other LULC classes in semi-arid environments. The results can be used for guiding and upscaling biological control options for managing mango insect pests such as the devastating invasive fruit fly Bactrocera dorsalis (Hendel) (Diptera: Tephritidae).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".