Alteration assemblage characterization using machine learning applied to high-resolution drill-core images, hyperspectral data and geochemistry
Bibliographic record
Abstract
Integration of multiple data types is beneficial for prediction of geological characteristics. From the perspective that geochemistry characterizes the composition of a rock mass, hyperspectral data characterizes alteration mineralogy and image feature extraction characterizes texture, most geological classifications would be well-informed by the combination of these three features. The process of meaningfully integrating distinctly sourced datasets and producing scale-relevant predictions for geological classifications involves several steps. We demonstrate a workflow to comprehensively structure and integrate these three feature families, refine training data, predict alteration classes and mitigate noise derived from scale mismatch in output predictions. The dataset, compiled from the Josemaria porphyry copper–gold deposit in Argentina, is comprised of more than 14 000 intervals of approximately 2 m, taken from 36 drillholes, where geochemistry was merged with hyperspectral mineralogy represented as tabular pixel abundances, and textural metrics extracted from core imagery, structured into the geochemical interval. Feature engineering and principal component analysis provided insights into the behaviour of the ore system during intermediate steps, as well as providing uncorrelated feature inputs for a random forest predictor. Training data were refined by producing an initial prediction, thresholding the predictions to >70% dominant class probability and using those (high-probability) samples to produce a final model encoding better constrained separation between alteration assemblages. Prediction using the final model returned an accuracy of 82.5%, as a function of model discrepancy combined with logging ambiguity and a scale mismatch between generalized logged intervals and much more granular (2 m) feature inputs. Noise reduction and generalization to the desired resolution of output was achieved by applying the multiscale multivariate continuous wavelet transform tessellation method to class membership probabilities. Ultimately, a large database of logged drill core was homogenized using empirical methodologies. The described workflow is adaptable to distinct scenarios with some modification and is apt for integrating multiple input feature types and using them to systematically define geological classifications in drillhole data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".