Bibliographic record
Abstract
With the increasing complexity of network environments, evaluating security risks can be a challenge.Quantitatively scoring computer systems or networks based on specific threats or concerns can facilitate the comparison of their security levels.Such scoring needs considering multiple security risk factors and multiple systems to enable meaningful comparisons, determining whether one system or network is relatively more secure than another.Our research aims to address this challenge by developing a comprehensive risk score aggregation methodology for evaluating security in modern network environments.In this thesis, we present two approaches to risk score aggregation: multi-factor aggregation for systems and multi-target aggregation for networks.In the first work, we propose an aggregation approach that combines machine measurements with human decision making for aggregating well-established individual risk factors into a risk score for a system.Specifically, we modify the Analytic Hierarchy Process (AHP) to facilitate group decision-making among selected "experts" who can derive the weights of individual risk factors for aggregation.We showcase the feasibility of our approach by selecting several common metrics to measure the target systems in our testbed and conducting an AHP survey with seventeen experts.The resulting overall aggregated scores for the target systems demonstrate how our approach enables the comparison of the overall security between those systems.By considering cloud-oriented settings, we also showcase how this approach i can be applicable to today's virtualized environments.The existing state-of-the-art approaches for aggregating multiple systems only consider one target in a network.The aggregation of risks scores of multiple targets is still yet to be explored.To this end, we propose another aggregation approach that uses the well-established attack paths as inter-system influences, treating each system as a potential target, to calculate cumulative attack probabilities for each system.We further aggregate these probabilities (i.e., the risk scores) into a single risk score for the entire network.To evaluate this approach, we consider several typical network types, including virtually hosted telecom networks such as 5G core and 5G Mobile Access Edge Computing (MEC) along with two additional network types.We employed the dataset collected from a real research data center.The results show that our approach can be generalized and applied to assess the security of various type of network.By considering multiple risk factors and interconnected entities/systems in network environments, we believe that our proposed methodologies for risk score aggregation can provide valuable insights and contribute to the research on risk assessments in modern IT environments.First and foremost, I would like to express my sincere gratitude to my thesis supervisor, Dr. Lianying Zhao for his constant support, invaluable feedback and patience throughout the entire research process.Without his guidance and trust, I can confidently say that I would not have reached this point in my academic career.I am also deeply thankful to our research manager, Makan Pourzandi for engaging in discussions, sharing insights and giving encouragement.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.012 | 0.041 |
| Meta-epidemiology (narrow) | 0.002 | 0.000 |
| Meta-epidemiology (broad) | 0.002 | 0.001 |
| Bibliometrics | 0.007 | 0.006 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.004 | 0.003 |
| Open science | 0.001 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".