Enhancing soil organic carbon estimation with generative AI and Nix color sensor
Notice bibliographique
Résumé
Soil organic carbon (SOC) is a key indicator of soil health, yet conventional laboratory assays are labor-intensive and costly. This study investigates a rapid and low-cost alternative by using a handheld Nix Spectro 2 Color Sensor, which captured high-resolution color data from air-dried soil samples. These color parameters were used to predict SOC with four data-driven prediction engines: Random Forest (RF), Gradient Boosting Regression (GBR), Extreme Gradient Boosting (XGBoost), and an Artificial Neural Network (ANN) and further strengthened them with synthetic data augmentation techniques. A total of 641 surface soil samples collected from six districts in West Bengal, India, were divided into 70% calibration and 30% validation subsets. Synthetic samples were produced using a combination of generative artificial intelligence (AI) techniques [generative adversarial networks (GANs) and Gaussian mixture models (GMM)] and non-parametric/statistical data augmentation methods [k-nearest neighbors (KNN) and bootstrapping] to fill critical gaps in the SOC range (3-14%). Among the baseline models using raw Nix color data, RF achieved the best validation accuracy (R² = 0.71, RMSE = 0.93%). After augmenting the calibration set with 44 GMM-generated samples (3-7% SOC), RF performance rose to R² = 0.77 and RMSE = 0.84%, while bias dropped and coverage across the SOC distribution improved markedly. The incorporation of synthetic data mitigated model bias and enhanced predictive accuracy despite Levene's test revealing significant variance differences between calibration and validation datasets. The enhanced generalization of the model was attributed to better coverage of the SOC distribution, reducing underrepresented gaps in the dataset. The study highlighted the potential of AI-driven soil monitoring techniques in precision agriculture, demonstrating that integrating the Nix color sensor with synthetic data augmentation, provides a rapid and cost-effective solution for on-site soil assessments. Future research should expand these methodologies to multi-parameter soil assessments, digital soil mapping, and broader applications in sustainable soil management and climate change mitigation.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction distillée sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Apprise à partir de 10 348 étiquettes directes de Codex et de 10 348 étiquettes directes de Gemma. Le mode candidate est l'union des têtes enseignantes seuillées; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont ni des étiquettes humaines ni des étiquettes directes de modèles de pointe.
Scores Codex et Gemma par catégorie
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,000 | 0,000 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,000 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,000 |
| Communication savante | 0,000 | 0,000 |
| Science ouverte | 0,000 | 0,000 |
| Intégrité de la recherche | 0,000 | 0,000 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,000 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule tête enseignante, pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».