A balanced mineral prospectivity model of Canadian magmatic Ni (± Cu ± Co ± PGE) sulphide mineral systems using conditional variational autoencoders
Bibliographic record
Abstract
With the increasing demand for raw materials, innovative exploration techniques are needed to discover large mineral deposits that are accessible from the surface. In recent years, various supervised machine learning techniques have proven effective for mineral prospectivity modelling (MPM). However, the successful application of these techniques has been limited due to the scarcity of known mineral deposits compared to barren regions, which leads to a model imbalance favouring the latter. We address the data imbalance challenge in MPM by proposing a novel generative modelling approach using a conditional variational autoencoder (CVAE). We compare the proposed method with two other data balancing techniques, namely the synthetic minority oversampling technique and class weighting. Furthermore, the efficacy of the balancing strategies is evaluated for three MPM classification methods, including extreme gradient boosting machines (XGBM), random forests, and multilayer perceptrons . We implement and test the approaches by modelling the prospectivity of magmatic Ni (±Cu ±Co ±Platinum group elements) sulphide mineral systems for the Canadian landmass. With an area under the success rate curve of 0.95 for a spatially distinct testing data set, we observe that a combination of the proposed CVAE framework with the XGBM classification model surpasses the other methods. Furthermore, the geographical representation of our XGBM-CVAE model demonstrates a strong association with known Ni mineral occurrences in Canada, along with new prospective regions in underexplored areas.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.001 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".