Rapid retrieval and classification of passive-source body wave events using a convolutional self-attention encoder: Application to gas storage monitoring
Bibliographic record
Abstract
ABSTRACT Passive-source seismic interferometry (SI) demonstrates significant potential in geophysical monitoring owing to its low cost and nondestructive attributes. However, uneven noise source distributions, pervasive surface waves, and other sources of noise often obscure body wave reflections, leading to very low signal-to-noise ratios in reconstructed signals. Additionally, long-term data acquisition generates large, high-noise data sets, increasing processing costs. Thus, rapidly retrieving body wave-dominated segments is essential for preprocessing. In this study, a convolutional self-attention encoder (CSE) was introduced to address these challenges. This method adopted multimodal inputs consisting of noise segments and their corresponding frequency attributes, enhancing overall performance. An attention mechanism was integrated to construct a binary classifier based on the CSE, enhancing contextual understanding. These steps helped the CSE to better learn features from the data sets, thereby improving generalization ability and prediction accuracy. Accordingly, the network was updated through incremental learning on small data sets, enabling continuous prediction and gradual adaptation to new data. This enabled us to rapidly generate large, high-quality training data sets and to train models with high performance. To further improve prediction data quality, the intermediate features extracted from the deep convolutional layers corresponding to the retrieved body wave events were extracted, and these features were clustered to perform quality classification of the retrieved data. Synthetic simulation data and field data from Northwest China validated that the proposed method retrieves and classifies body wave reflections with high accuracy and efficiency, providing robust data for subsequent monitoring and enhancing imaging results. The proposed method provides efficient and reliable data preprocessing for the application of passive-source SI in energy transition scenarios, such as subsurface fluid migration monitoring and carbon dioxide geological sequestration, and facilitates the development of low-carbon exploration and green transformation in the energy industry.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".