EURATOM METIS D4.1 Seismic source characterizations methodologies and applications
Bibliographic record
Abstract
Probabilistic seismic hazard assessment considers the occurrence of earthquakes as Poissonian, each even beingindependent of the others. As earthquakes do not occur as isolated events but for clusters of foreshocks, mainshock andaftershocks, it is necessary to identify and remove the earthquakes that are other than mainshocks from the catalog beforecalculating the annual rate of events used for the characterization of the seismic activity of the seismogenic sources of thehazard model. The aim of this study is to find a methodology to obtain a Poissonian declustered catalog while removing as fewearthquakes as possible. To do so, we developed a methodology that allows to test the Poissonian nature of a declusteredearthquake catalog and find the declustering algorithm that keeps the largest number of events while still performing well onthe Poissonian test. Building a hazard model requires to perform statistic on the number of events per units of time and theirspatial repartition, therefore, the Poissonian test was developed to reflect these needs. By comparing the inter-eventspatio-temporal distances between the events of the tested catalog to the ones from simulated Poissonian catalogs sharingsimilar properties, we can score the statistical similarities between the tested and the simulated catalogs. We apply thismethodology on the catalog for central Italy after being declustered with a range of different published algorithms and anadditional algorithm that we propose in this study, and we perform a seismic hazard study for two locations in Central Italy toquantify the impact of the declustering algorithm on the seismic hazard level estimate.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.002 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.018 | 0.010 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".