Identifying serpentine minerals by their chemical compositions with machine learning
Bibliographic record
Abstract
Abstract The three main serpentine minerals, chrysotile, lizardite, and antigorite, form in various geological settings and have different chemical compositions and rheological properties. The accurate identification of serpentine minerals is thus of fundamental importance to understanding global geochemical cycles and the tectonic evolution of serpentine-bearing rocks. However, it is challenging to distinguish specific serpentine species solely based on geochemical data obtained by traditional analytical techniques. Here, we apply machine learning approaches to classify serpentine minerals based on their chemical compositions alone. Using the Extreme Gradient Boosting (XGBoost) algorithm, we trained a classifier model (overall accuracy of 87.2%) that is capable of distinguishing between low-temperature (chrysotile and lizardite) and high-temperature (antigorite) serpentines mainly based on their SiO2, NiO, and Al2O3 contents. We also utilized a k-means model to demonstrate that the tectonic environment in which serpentine minerals form correlates with their chemical compositions. Our results obtained by combining these classification and clustering models imply the increase of Al2O3 and SiO2 contents and the decrease of NiO content during the transformation from low-to high-temperature serpentine (i.e., lizardite and chrysotile to antigorite) under greenschist–blueschist conditions. These correlations can be used to constrain mass transfer and the surrounding environments during the subduction of hydrated oceanic crust.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".