MétaCan
Menu
Back to cohort
Record W4408425925 · doi:10.5194/egusphere-egu25-7718

Dataset preparation for Resolving Apparent Source Time Functions (ASTF) and Evalutions Using Basic ML-models

2025· preprint· en· W4408425925 on OpenAlexaffabout
Runcheng Pang, Hongyu Yu, Ge Li, Haoran Meng, Zaiwang Liu, Cheng Su, Deli Zha, Wanli Tian

Bibliographic record

Venuenot available
Typepreprint
Languageen
FieldPhysics and Astronomy
TopicModel Reduction and Neural Networks
Canadian institutionsMila - Quebec Artificial Intelligence Institute
Fundersnot available
KeywordsComputer science

Abstract

fetched live from OpenAlex

Monitoring induced seismic harzards during the fluid injections has been a significant challenge for geo-energy development. Conventional approaches, such as the (adaptive) traffic light protocol and prediction methods based on statistical and machine learning regressions, often yield limited accuracy due to diverse geological conditions across regions. A promising direction lies in developing precursors to monitor fault reactivation more effectively.Recent seismological studies have shown that aseismic slip loading is more prevalent than previously thought during fault activation induced by fluid injections (Yu et al., 2021a, b; Eyre et al., 2019, 2022). Slow earthquake signals, like Earthquakes characterized by Hybrid-frequency Waveforms (EHW), occur during the transition from fault creep to brittle rupture induced by fluid injection (Guglielmi et al., 2015). These signals are potential indicators for aseismic slip loading and fault reactivation. However, their longer rupture durations distinguish them from typical induced earthquakes, rendering classic source analysis methods ineffective for real-time monitoring.To address this limitation, we propose ASTF-Net, a machine learning (ML) model designed to predict Apperant Source Time Function (ASTF) by deconvoluting Empirical Green's Function (EGF) waveform from target waveform in time domain. This approach provides reliable real-time estimates of source durations, to identify slow earthquake signals, specifically EHW, and offers a valuable tool for fault activation monitoring. A robust and well-sampled dataset is therefore crucial for the model's performance.In this study, we present a dataset designed for developing a single-channel version of ASTF-Net and evaluate its effectiveness using basic ML-models. The dataset consists of three parts: ASTFs, EGFs, and target seismic waveforms. Synthetic ASTFs are generated using kinematic forward modeling with an elliptical rupture model to simulate earthquake events with magnitudes ranging from Mw 3.0 to 4.5 and stress drops between 5 and 20 MPa. These random ASTFs are calculated under various ray paths. We collect EGFs from hydraulic fracturing-induced earthquakes (M1.5-2.5) in the Southern Montney Play, western Canada, recorded by a network of 40 nodal/broadband seismic stations between 2017 and 2020. Synthetic target waveforms are then created by convolving ASTFs with corresponding EGFs. The dataset’s inputs consist of single-channel synthetic seismic waveforms and their corresponding EGFs, with ASTFs as outputs. To assess the generalization and robustness of the model, records from different stations are divided into training, validation, and four test sets with varying difficulty levels based on geographical locations and event counts. Notably, Test level 4 is human analysis results reported by Roth et al. (2022).We evaluate model performance using basic ML architectures, including MLP, CNN, VGG and U-Net. Performance metrics include the Correlation Coefficient (CC) between predictions and labeled ASTFs, and the relative error in apparent source duration. CNN emerges as the most promising candidate for the further optimization, achieving the following CC > 0.9 across test levels: 87.9% (Level 1), 85.9% (Level 2), and 80.3% (Level 3). The percentages of relative error below 10% are 59.3%, 59.3%, and 51.2%, respectively, for the three levels.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Methods · Consensus signal: none
Teacher disagreement score0.960
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0000.000
Science and technology studies0.0000.000
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0000.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.060
GPT teacher head0.335
Teacher spread0.275 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designSimulation or modeling
Domainnot available
GenreMethods

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2025
Admission routes2
Has abstractyes

Explore more

Same topicModel Reduction and Neural NetworksFrench-language works237,207