Assessment of Soil Suitability Using Machine Learning in Arid and Semi-Arid Regions
Bibliographic record
Abstract
Increasing agricultural production is a major concern that aims to increase income, reduce hunger, and improve other measures of well-being. Recently, the prediction of soil-suitability has become a primary topic of rising concern among academics, policymakers, and socio-economic analysts to assess dynamics of the agricultural production. This work aims to use physico-chemical and remotely sensed phenological parameters to produce soil-suitability maps (SSM) based on Machine Learning (ML) Algorithms in a semi-arid and arid region. Towards this goal an inventory of 238 suitability points has been carried out in addition to14 physico-chemical and 4 phenological parameters that have been used as inputs of machine-learning approaches which are five MLA prediction, namely RF, XgbTree, ANN, KNN and SVM. The results showed that phenological parameters were found to be the most influential in soil-suitability prediction. The validation of the Receiver Operating Characteristics (ROC) curve approach indicates an area under the curve and an AUC of more than 0.82 for all models. The best results were obtained using the XgbTree with an AUC = 0.97 in comparison to other MLA. Our findings demonstrate an excellent ability for ML models to predict the soil-suitability using physico-chemical and phenological parameters. The approach developed to map the soil-suitability is a valuable tool for sustainable agricultural development, and it can play an effective role in ensuring food security and conducting a land agriculture assessment.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".