MétaCan
Menu
Back to cohort
Record W4292014083 · doi:10.1002/saj2.20469

Machine learning models for predicting soil particle size fractions from routine soil analyses in Quebec

2022· article· en· W4292014083 on OpenAlexafffundabout
Gaëtan Martinelli, Marc‐Olivier Gasser

Bibliographic record

VenueSoil Science Society of America Journal · 2022
Typearticle
Languageen
FieldEnvironmental Science
TopicSoil Geostatistics and Mapping
Canadian institutionsInstitut de Recherche et de Développement en Agroenvironnement
FundersOntario Ministry of Agriculture, Food and Rural AffairsMinistère de l'Agriculture, des Pêcheries et de l'Alimentation
KeywordsSoil textureSiltSoil scienceParticle sizeSoil testCation-exchange capacityParticle-size distributionLinear regressionUSDA soil taxonomyRandom forestSoil waterBulk densityParticle (ecology)Soil typeSoil classificationEnvironmental scienceMathematicsStatisticsMachine learningGeologyComputer science

Abstract

fetched live from OpenAlex

Abstract Soil texture and particle size distribution are important soil properties for understanding and interpreting the multiple processes and interactions in soil agrosystems. However, soil particle size analysis is rarely included in routine soil laboratory analyses due to cost but can be derived from other soil physico‐chemical properties and from more powerful regressions using machine learning techniques. The performance of multiple linear regression and four machine learning algorithms (MLAs)—K‐nearest neighbor (KNN), Random Forest, extra‐gradient boosting (XGBoost), and multilayer neural network (NeuralNetwork)—was compared to predict particle size fractions of clay, sand, and silt using routine analyses of soil pH, cation exchange capacity, Mehlich‐3 extracted elements, and density of dried sieved soil from 8,364 soil samples distributed across Quebec. Particle size fractions were predicted as compositional data using isometric log ratio transformation. The XGBoost model performed best, with RMSE values of 7, 10, and 12% and R 2 values of .77, .57, and .73 for the prediction of clay, silt, and sand fractions, respectively. Feature importance classification varied from one model to another, but the best predictors remained the same regardless of the model used. Sieved soil density measured in the laboratory and Mehlich‐3 Mg, K, Fe, and Zn ranked as better predictors of particle size fractions. Predicting soil particle sizes using routine soil physico‐chemical properties and MLAs appears an option for incomplete legacy datasets and a promising economic alternative to current methods carried out in commercial laboratories.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesScience and technology studies
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.103
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0010.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.001
Science and technology studies0.0020.001
Scholarly communication0.0000.001
Open science0.0000.000
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.029
GPT teacher head0.286
Teacher spread0.257 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations12
Published2022
Admission routes3
Has abstractyes

Explore more

Same venueSoil Science Society of America JournalSame topicSoil Geostatistics and MappingFrench-language works237,207