Mapping Crustal Vp/Vs in North America With a Machine Learning Approach
Bibliographic record
Abstract
Abstract Vp/Vs (Poisson's ratio) provides critical information for constraining the bulk crustal composition, stress state, and tectonic evolution of the Earth. The receiver function technique has been extensively utilized to constrain the crustal Vp/Vs, yet the reliability of measurements can be affected by complex structures and uneven distribution of seismic stations. Consequently, the interpolated Vp/Vs maps can often be biased by unreliable observations, especially in data‐sparse regions. We tackle these issues by proposing a machine learning model that integrates multiple geophysical data sets to estimate Vp/Vs, leveraging the physical and structural properties of the crust. We train the model by compiling an extensive data set of global Vp/Vs measurements at 13,314 seismic stations and employ XGBoost to map Vp/Vs with other key crustal properties. Experiments using data from the (a) United States and (b) United States and Canada demonstrate superior prediction accuracy, achieving an overall value of 0.84 in both cases. Feature importance analysis indicates that crustal tectonic type, geographic coordinates, mid‐crust shear‐wave velocity, and crustal thickness primarily capture Vp/Vs variations, together explaining over 70% of reduction in the normalized root‐mean‐square error. The inclusion of other features further refines small‐scale Vp/Vs variation. Compared to cubic and Kriging interpolations, the predicted Vp/Vs map from machine learning exhibits less local extremes and a better alignment with the first‐order crustal structure across the continent. This study highlights the capability of machine learning to uncover complex geophysical relationships for reliable Vp/Vs estimates and its potential to constrain crustal composition at a continental scale.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".