A balanced mineral prospectivity model of Canadian magmatic Ni (± Cu ± Co ± PGE) sulphide mineral systems using conditional variational autoencoders
Bibliographic record
Abstract
With the increasing demand for raw materials, innovative exploration techniques are needed to discover large mineral deposits that are accessible from the surface. In recent years, various supervised machine learning techniques have proven effective for mineral prospectivity modelling (MPM). However, the successful application of these techniques has been limited due to the scarcity of known mineral deposits compared to barren regions, which leads to a model imbalance favouring the latter. We address the data imbalance challenge in MPM by proposing a novel generative modelling approach using a conditional variational autoencoder (CVAE). We compare the proposed method with two other data balancing techniques, namely the synthetic minority oversampling technique and class weighting. Furthermore, the efficacy of the balancing strategies is evaluated for three MPM classification methods, including extreme gradient boosting machines (XGBM), random forests, and multilayer perceptrons . We implement and test the approaches by modelling the prospectivity of magmatic Ni (±Cu ±Co ±Platinum group elements) sulphide mineral systems for the Canadian landmass. With an area under the success rate curve of 0.95 for a spatially distinct testing data set, we observe that a combination of the proposed CVAE framework with the XGBM classification model surpasses the other methods. Furthermore, the geographical representation of our XGBM-CVAE model demonstrates a strong association with known Ni mineral occurrences in Canada, along with new prospective regions in underexplored areas.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".