Machine learning applied to apatite compositions for determining mineralization potential
Bibliographic record
Abstract
Abstract Apatite major and trace element chemistry is a widely used tracer of mineralization as it sensitively records the characteristics of the magmatic-hydrothermal system at the time of its crystallization. Previous studies have proposed useful indicators and binary discrimination diagrams to distinguish between apatites from mineralized and unmineralized rocks; however, their efficiency has been found to be somewhat limited in other systems and larger-scale data sets. This work applied a machine learning (ML) method to classify the chemical compositions of apatites from both fertile and barren rocks, aiming to help determine the mineralization potential of an unknown system. Approximately 13 328 apatite compositional analyses were compiled and labeled from 241 locations in 27 countries worldwide, and three apatite geochemical data sets were established for XGBoost ML model training. The classification results suggest that the developed models (accuracy: 0.851–0.992; F1 score: 0.839–0.993) are much more accurate and efficient than conventional methods (accuracy: 0.242–0.553). Feature importance analysis of the models demonstrates that Cl, F, S, V, Sr/Y, V/Y, Eu*, (La/Yb)N, and La/Sm are important variables in apatite that discriminate fertile and barren host rocks and indicates that V/Y and Cl/F ratios and the S content, in particular, are crucial parameters to discriminating metal enrichment and mineralization potential. This study suggests that ML is a robust tool for processing high-dimensional geochemical data and presents a novel approach that can be applied to mineral exploration.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".