Vulnerability scoring metric of CVSS needs to be adjusted per each product: our analysis on Linux and Apache
Bibliographic record
Abstract
To safeguard software products against security risks, it is imperative for organizations to prioritize and rank vulnerabilities in a systematic manner. The Common Vulnerability Scoring System stands as a widely embraced open standard for assessing and classifying vulnerabilities in computer systems. The precision of the CVSS methodology is of paramount importance, enabling security experts to appropriately prioritize vulnerabilities and make well-informed decisions in response to security threats. Despite the utilization of a standardized approach in the computation of the CVSS formula, it is noteworthy that certain shortcomings persist, which have not yet been adequately addressed in previous research. This paper presents a fresh perspective on these aforementioned limitations and presents innovative remedies aimed at enhancing the precision of vulnerability severity scores. The absence of these enhancements precludes the attainment of desirable outcomes. Our empirical investigation, which involves an in-depth analysis of CVSS outcomes for both Linux and Apache products, emphasizes the need to tailor the CVSS formula individually for each product to ensure the accurate determination of vulnerability severity scores. In our study, we reveal two distinctive insights that have the potential to increase the effectiveness of the CVSS methodology. The first insight revolves around the consideration of module frequency within a product, coupled with vigilant monitoring of vulnerability occurrences within those specific modules. This approach allows the calculation of severity scores to deviate for modules characterized by elevated vulnerability levels, as opposed to other modules governed by identical CVSS parameters. The second insight pertains to the nuanced weighting of parameter values within the foundational metrics of the CVSS methodology. Our evaluation findings emphasize that the majority of attacks on the Linux product do not require any elevated privileges, whereas for the Apache product, a minimum threshold of permission is a prerequisite. Thus, the formulation of this weighting mechanism should take inspiration from the historical behavior of past vulnerabilities within the product. In conclusion, the accurate assessment and prioritization of vulnerabilities are critical in fortifying software products against potential security breaches. The Common Vulnerability Scoring System serves as an established framework in this endeavor, yet its imperfections persist. Our study aims to address these shortcomings and introduce novel perspectives that can amplify the precision of vulnerability severity scores. By tailoring the CVSS formula on a product-specific basis and incorporating insights derived from module frequency and historical vulnerability behaviors, the proposed enhancements aim to empower security experts to make more informed and effective decisions in mitigating software vulnerabilities.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.007 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.002 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".