Generation and Evaluation of Synthetic Low‐Magnitude Earthquake Data Using Auxiliary Classifier GAN
Bibliographic record
Abstract
Abstract Low‐magnitude earthquakes occur far more frequently than major quakes and often go unnoticed by the public. These tremors rarely cause any damage, yet they play an important role in advancing our understanding of Earth's seismicity. Accurate detection of low‐magnitude earthquakes is crucial to develop complete earthquake catalogs and improve seismic hazard forecasting models. However, conventional detection algorithms such as the short‐time‐average/long‐time‐average (STA/LTA) method struggle to identify these events because of their inherently low signal‐to‐noise ratio (SNR). Additionally, lack of labeled waveforms for low‐magnitude earthquakes further complicates the training of effective deep‐learning models. In this study, we use an Auxiliary Classifier Generative Adversarial Network (AC‐GAN) to produce synthetic yet realistic three‐component waveforms of low‐magnitude earthquakes. The AC‐GAN is trained on fixed‐length (60‐s) waveform segments conditioned by predefined SNR classes. All selected events have magnitudes lower than 3 and are categorized into 10 distinct SNR classes. Our results indicate that the AC‐GAN model generates realistic three‐component waveforms that effectively capture essential characteristics of real seismic signals. To evaluate the quality of these synthetic waveforms, we employ both quantitative and qualitative assessments. Quantitative analysis using Pearson's correlation coefficient yield relatively low correlations (ranging from 0.01 to 0.04); however, correlation values noticeably improve as SNR increases. Qualitatively, a user‐based visual inspection experiment demonstrate remarkable similarities in general seismic features between the synthetic and authentic waveforms. We also test their effectiveness for data augmentation in binary deep‐learning classifier designed for detecting low‐magnitude earthquakes. Our result show improved classification performance with the addition of synthetic data.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.000 | 0.001 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".