Predicting soil organic matter and soil moisture content from digital camera images: comparison of regression and machine learning approaches
Bibliographic record
Abstract
Appropriate soil management maintains and improves the health of the entire ecosystem. Soil appropriate administration necessitates proper characterization of its properties including soil organic matter (SOM) and soil moisture content (SMC). Image-based soil characterization has shown strong potential in comparison with traditional methods. This study compared the performance of 22 different supervised regression and machine learning algorithms, including support vector machines (SVMs), Gaussian process regression (GPR) models, ensembles of trees, and artificial neural network (ANN), in predicting SOM and SMC from soil images taken with a digital camera in the laboratory setting. A total of 22 image parameters were extracted and used as predictor variables in the models in two steps. First models were developed using all 22 extracted features and then using a subset of six best features for both SOM and SMC. Saturation index (redness index) was the most important variable for SOM prediction, and contrast (median S) for SMC prediction, respectively. The color and textural parameters demonstrated a high correlation with both SOM and SMC. Results revealed a satisfactory agreement between the image parameters and the laboratory-measured SOM ( R 2 and root mean square error (RMSE) of 0.74 and 9.80% using cubist) and SMC ( R 2 and RMSE of 0.86 and 8.79% using random forest) for the validation data set using six predictor variables. Overall, GPR models and tree models (cubist, RF, and boosted trees) best captured and explained the nonlinear relationships between SOM, SMC, and image parameters for this study.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".