Combining Machine Learning and Geophysical Inversion for Applied Geophysics
Bibliographic record
Abstract
Machine learning and geophysical inversion both represent ways that the applied geophysicist might gain knowledge from field observations and remote sensed data. The two approaches represent contrasting philosophies based respectively on statistics and physics. Both potentially add insights which might help constrain 3D geology by geophysical means. Machine learning uses patterns in data to provide statistically controlled predictions, e.g. of lithology. In contrast, geophysical inversion relies on modelling the physical response of 3D geological block geometry in a deterministic manner. Although both approaches are widely used, it is not currently commonplace in applied geosciences to make use of a combined approach.We present an example which aims to refine the 3D geology in a prospective region of west Tasmania. Although the region is geologically well-mapped, thick vegetation and significant topography present a challenging set of conditions under which to refine the lithology and block geometry to a level of detail which will support the next generation of exploration. We use multiple layers of remote sensed geophysical data to provide probabilistic information on near-surface lithology extent using the Random Forests classifier. We show how the statistical, robust, output from the machine learning exercise can be used to guide the construction of improved volume geometry within a 3D GOCAD geological and geophysical modelling environment. This enables better constraints to be supplied to the geophysical inversion with resulting improvements in the detail of the 3D geology.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.010 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.002 |
| Scholarly communication | 0.003 | 0.002 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.003 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".