MétaCan
Menu
Retour à la cohorte
Enregistrement W3205087287 · doi:10.25959/100.00037833

Machine learning for mineral exploration: prediction and quantified uncertainty at multiple exploration stages

2021· dissertation· en· W3205087287 sur OpenAlexaboutno aff
S Kuhn

Notice bibliographique

RevueUTAS Research Repository · 2021
Typedissertation
Langueen
DomaineComputer Science
ThématiqueGeochemistry and Geologic Mapping
Établissements canadiensnon disponible
Organismes subventionnairesnon disponible
Mots-clésMineral explorationVariety (cybernetics)Context (archaeology)GeologistMachine learningComputer scienceArtificial intelligenceGround truthData miningData scienceGeologyGeophysics

Résumé

récupéré en direct d'OpenAlex

Machine learning describes an array of computational and nested statistical methods whereby a computer can 'learn' and subsequently make predictions or identify patterns in data. With the increasing volume and variety of numerical data in the geosciences, and widespread availability of the needed computing power, machine learning techniques are a logical addition to the numerous possible approaches that can be applied to the search for ore deposits. The three core research chapters in this thesis develop the application of machine learning in the context of mineral exploration. Emphasis is placed on the Random Forests algorithm for mapping lithology in a range of settings and at a variety of stages in the exploration process. Information entropy is used to assist both in assessing and communicating any complex combinations, and potential inaccuracy, of classification results. Through the thesis, methods are employed with future practical usage in mind, such that machine learning may be used by the geologist (as domain expert) in an objective manner. The first of these core studies uses the Random Forests algorithm to re-classify the solid geology lithology map of the Heron South project, located in the Eastern Goldfields ofWestern Australia. This study uses geophysical and remote sensing data, in the absence of geochemical samples and geological ground truthing with most of the project under transported cover. This is characteristic of an early stage, reconnaissance exploration project. A sparse training sample of 1.6 percent of the total area, is taken as training data, allowing much of the areas geology the freedom to be reclassified. This study demonstrates that Random Forests, with proper consideration given to sampling and training data selection, can be used effectively to produce or improve geological mapping in little-explored areas. Information entropy is shown to be valuable in predicting where classification was likely to be inaccurate or a region highly complex. The second core study uses Random Forests to produce a solid geology map of the Kliyul porphyry prospect of British Columbia, Canada, using a fusion of available geophysical and geochemical data, typical of a greenfields stage exploration project. Soil and rock chip sample sites were taken as training data, used to classify the remainder of the project area. Assessment of the probability distributions produced using the Random Forests algorithm enabled regions with an elevated probabilityof intrusions (a key indicator lithology) to be mapped, even where not observed in training data. The results of this study highlight the value of a soft, ensemble classifier such as Random Forests, and the value to be gained from an assessment of the spatial distribution of class probabilities as opposed to viewing a final map as a solution in isolation. In the third and final core study, a range of training data sampling paradigms are tested in a data rich area located in the Domes region of the Central African Copper belt hosting the Sentinel (Ni) and Enterprise (Cu) deposits. This study simulates early and advanced stage exploration project maturity in incorporating a priori geological Information. It culminates in the use of Random Forests to undertake an objective audit of the present company geological map. Further to this, unsupervised clustering is used in the production of a geological map in the absence of training or constraint through identifying the natural grouping of data. The results of these studies highlight the importance of proper sample balancing and explore the repercussions of limited and/or non-representative training data. The use of the information entropy proxy is developed to identify where a classification may depart from the domain represented by training data. The ranking of input data that is performed in association with the Random Forests classification can be used to improve clustering results through optimising dataset selection. Through the three core research chapters, a set of practical considerations and recommendations for explorers are provided. It is demonstrated that Random Forests can provide an objective audit and subsequent refinement of a pre-existing geological map. The expression of uncertainty using information entropy, and the assessment of class probabilities, can be used to appraise the results from the machine learning analyses. This includes validation in the case of complex outcome combinations, and generation of new insights. Ranking of input datasets via Random Forests can enhance understanding of data and improve both Random Forests classification results and improve clustering. With the proper selection of appropriate datasets, clustering (for example immobile trace elements) and scaling can indeed produce results that correspond well with lithology. Studies presented in this thesis use data from current/active exploration projects and methods are distilled to streamlined workflows using industry standard software and data formats. In summary, these methods, previously the domain of computer and data scientists, are now developed to be more widely accessible to mineral explorers.

Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.

Comment cette classification a été obtenuedéplier

Prédiction machine sur la base complète

Imitation des enseignants

Ni prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.

score de la tête « metaresearch » (Codex)0,010
score de la tête « metaresearch » (Gemma)0,032
Version: metacan-v3-hybrid-931329e0061cStatut de validation: machine_predicted_unvalidated
Catégories candidatesaucune
Catégories consensuellesaucune
DomaineSignal candidat: aucune · Signal consensuel: aucune
Devis d'étudeSignal candidat: Simulation ou modélisation · Signal consensuel: Simulation ou modélisation
GenreSignal candidat: Empirique · Signal consensuel: aucune
Score de désaccord entre enseignants0,010
Score d'incertitude au seuil0,052

Scores du classifieur distillé par catégorie (deux têtes)

CatégorieCodexGemma
Métarecherche0,0100,032
Méta-épidémiologie (sens strict)0,0010,001
Méta-épidémiologie (sens large)0,0020,001
Bibliométrie0,0020,002
Études des sciences et des technologies0,0010,004
Communication savante0,0050,008
Science ouverte0,0020,004
Intégrité de la recherche0,0020,005
Charge utile insuffisante (le modèle a refusé de juger)0,0020,000

Scores machine (provisoires)

Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.

Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.

Tête enseignante Opus0,084
Tête enseignante GPT0,331
Écart entre enseignants0,247 · la distance entre les deux têtes enseignantes sur ce seul travail
Statut de validationscore_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découle

Classification

machine, non validée

Prédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.

Les modèles n’ont appliqué aucune catégorie : rien dans la taxonomie ne correspondait à ce travail.
Devis d'étudeSimulation ou modélisation
Domainenon disponible
GenreEmpirique

Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».

En bref

Citations1
Publié2021
Routes d'admission1
Résumé présentoui

Explorer davantage

Même revueUTAS Research RepositoryMême sujetGeochemistry and Geologic MappingTravaux en français237 207