Support Vector Machine Classification of Seismic Events in the Tianshan Orogenic Belt
Bibliographic record
Abstract
Abstract Discriminating between various types of seismic events is of significant scientific and societal importance. We use a machine learning method employing support vector machine (SVM) to classify tectonic earthquakes (TEs), quarry blasts (QBs), and induced earthquakes (IEs) among 30,181 1.5 < M L <2.9 seismic events that occurred in the Tianshan orogenic belt in China from 2009 to 2017. SVM classifiers are derived based on discriminant features of a training data set consisting of 1,400 TEs selected from the aftershock sequences of 18 M L ≥ 5.0 earthquakes, 2,881 QBs from repeating events occurring in those areas with a percentage of event daytime occurrence greater than 0.9, and 987 IEs from events in the known oil/gas fields and water reservoirs. The discriminant features include spectral amplitudes of observed P and S wave signals in a frequency range of 1–15 Hz normalized by the P spectrum and averaged over the entire seismic network, and an optional feature of the percentage of event daytime occurrence. Statistics analyses indicate that the accuracies of the SVM classifiers are 99.81% for TEs, 99.93% for QBs, and 99.62% for IEs. Our classification indicates that 37.57% of the seismic events are QBs occurring in possible mine areas and appearing mostly as clusters with a percentage of event daytime occurrence greater than 0.9, 50.12% are TEs occurring in various thrust faults in the Tianshan orogenic belt, and 12.31% are IEs or shallow tectonic earthquakes occurring mostly as clusters near oil and gas fields and water reservoirs. We reevaluate b values in the region and obtain relatively uniform values for the classified TEs with most of them below 1.0, as opposed to a large range of values (0.5–2.7) when all the seismic events are used in the analysis.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".