CITI Fault Report Classification and Encoding for Vulnerability and Risk Assessment of Interconnected Infrastructures
Bibliographic record
Abstract
Eective functionalities of many of the critical infrastructures depend on Communication and Information Technology Infrastructure (CITI). As such, any fault in CITI can disrupt the operation of these infrastructures. Understanding the origin of these faults, their propagation pattern and their impact on other infrastructures can be very valuable for secure and reliable infrastructures design and operation. However, up to now there is no well-defined technique to comprehend these interinfrastructure fault scenarios. Public domain CITI fault reports can serve as a useful source to identify vulnerability patterns and impact of those vulnerabilities on other infrastructures. But, as most of these reports are unstructured description of fault events, this make their use limited and ineective for formal research. Until now, not much work was done to methodically classify and interpret these reports. However, such classification could give infrastructure research community huge benefit to explore this massive amount of open source information. In this paper, we propose a classification method and a report layout format, which will enable meaningful analysis of these fault reports and will enable selective query and filtering when kept in a database. We have demonstrated our method by classifying and analyzing some of those reports and have explained the results in the context of interdependency research.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.030 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.010 | 0.007 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.003 | 0.004 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".