Catastrophe Predictors From Ensemble Decision-Tree Learning of Wide-Area Severity Indices
Bibliographic record
Abstract
Catastrophe precursors are essential prerequisites for response-based remedial action schemes, at both the protective and the operator levels. In this paper, wide-area-severity indices (WASI) derived from PMU measurements serve as the basis for building fast catastrophe predictors using random-forest (RF) learning. Given the randomness in the ensemble of decision trees (DTs) stacked in the RF model, it can provide at the recall stage not only an early assessment of the stable/unstable status of an ongoing contingency but also a probability outcome which quantifies the confidence level of the decision. This methodology, which to the best of our knowledge is new to the dynamic security assessment (DSA) of power systems, is also very effective in evaluating the importance of and interaction among the various WASI input features. Our research unexpectedly showed that the ensemble of trees in the RF is very robust in the presence of small changes in the training data and generalize across widely different network dynamics. Thus, the same RF performed very well on a large database with more than 60 000 instances from a test system (10%) and an actual (90%) system combined. One such a general RF (with 210 trees) boosted the reliability of a 9-cycle catastrophe predictor to 99.9%, compared to only 70% when a single conventionally trained DT is used.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.007 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".