MétaCan
Menu
Back to cohort
Record W4226366790 · doi:10.1109/tdsc.2022.3170011

Classifier Calibration: With Application to Threat Scores in Cybersecurity

2022· article· en· W4226366790 on OpenAlexaff
Waleed A. Yousef, Issa Traoré, William Briguglio

Bibliographic record

VenueIEEE Transactions on Dependable and Secure Computing · 2022
Typearticle
Languageen
FieldComputer Science
TopicAdvanced Malware Detection Techniques
Canadian institutionsUniversity of Victoria
Fundersnot available
KeywordsClassifier (UML)Computer scienceLogistic regressionArtificial intelligenceCalibrationMachine learningBinary numberBinary classificationPattern recognition (psychology)Data miningStatisticsSupport vector machineMathematics

Abstract

fetched live from OpenAlex

This article explores the calibration of a classifier output score in binary classification problems. A calibrator is a function that maps the arbitrary classifier score, of a testing observation, onto [0,1] to provide an estimate for the posterior probability of belonging to one of the two classes. Calibration is important for two reasons; first, it provides a meaningful score, that is the posterior probability; second, it puts the scores of different classifiers on the same scale for comparable interpretation. The article presents three main contributions: (1) Introducing multi-score calibration, when more than one classifier provides a score for a single observation. (2) Introducing the exact analogy between two scenarios: (a) designing a classifier from a set of features, and (b) designing a calibrator, to generate a single calibrated score, from a set of scores of different classifiers. Hence, we propose expanding these classifiers’ scores to higher dimensions to boost the calibrator’s performance. (3) Conducting a massive simulation study, in the order of 24,000 experiments, that incorporates different configurations, in addition to experimenting on three real datasets from the cybersecurity domain. The results show that there is no overall winner among the different calibrators and different configurations. However, general advices for practitioners include the following: the Platt’s calibrator (J. Plattet al., 1999), a version of the logistic regression that decreases bias for a small sample size, has a very stable and acceptable performance among all experiments; our suggested multi-score calibration provides better performance than single score calibration in the majority of experiments, including the two real datasets. In addition, expanding the scores can help in some experiments.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.024
metaresearch head score (Gemma)0.129
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: none
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.024
Threshold uncertainty score0.128

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0240.129
Meta-epidemiology (narrow)0.0020.001
Meta-epidemiology (broad)0.0020.001
Bibliometrics0.0050.004
Science and technology studies0.0010.003
Scholarly communication0.0040.006
Open science0.0030.005
Research integrity0.0040.007
Insufficient payload (model declined to judge)0.0030.001

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.009
GPT teacher head0.236
Teacher spread0.227 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations8
Published2022
Admission routes1
Has abstractyes

Explore more

Same venueIEEE Transactions on Dependable and Secure ComputingSame topicAdvanced Malware Detection TechniquesFrench-language works237,207