Mineralogy from µXRF elemental maps: comparison of supervised and unsupervised methods
Bibliographic record
Abstract
This study evaluates the effectiveness of four methods for estimating mineralogy from micro X-ray fluorescence (µXRF) elemental maps: two optimized/supervised methods (linear programming and Random Forests) and two unsupervised methods (k-means clustering and Uniform Manifold Approximation and Projection (UMAP) combined with Hierarchical Density Based Spatial Clustering of Applications with Noise (HDBSCAN)). µXRF is a non-destructive, high-resolution imaging technique that provides spatially resolved elemental data and bridges the scale gap between micro- and macroscales in mineralogical studies. Analysing samples from porphyry copper deposits (representing a range of textures and mineralogical complexities), each method was assessed based on its ability to accurately represent the mineralogy and texture of each sample relative to a petrographic image. Compared to traditional petrographic assessment, Random Forests yielded the lowest normalized root mean squared error (NRMSE) values (0.09–0.38), while k-means (0.06–0.44), linear programming (0.07–0.44) and UMAP–HDBSCAN (0.18–0.64) showed more variable and generally higher error values across the samples. k-Means and linear programming produced results within minutes, Random Forests in ∼10 minutes and UMAP–HDBSCAN in ∼20 minutes. The results demonstrate that supervised methods, particularly Random Forests, deliver better mineralogical estimates, especially in complex mineral assemblages. However, unsupervised methods (particularly k-means) offer a faster initial assessment, making them valuable for preliminary analysis. Combining different methods using known mineral compositions or training datasets could enhance cluster-to-mineral correlation. The integration of supervised, optimization and unsupervised approaches for mineralogy from µXRF data can enhance the robustness and efficiency of mineralogical interpretations, adding significant value across mineral exploration and processing.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".