MétaCan
Menu
Back to cohort
Record W4402169893 · doi:10.32920/26883781

Stochastic-Based Hyperparameter Selection and Learnability Analysis in Supervised and Unsupervised Learning

2024· preprint· en· W4402169893 on OpenAlexaff
Mahdi Shamsi

Bibliographic record

Venuenot available
Typepreprint
Languageen
FieldComputer Science
TopicMachine Learning and Data Classification
Canadian institutionsToronto Metropolitan University
Fundersnot available
KeywordsLearnabilityHyperparameterMachine learningArtificial intelligenceComputer scienceSelection (genetic algorithm)Statistical learningUnsupervised learningSupervised learningArtificial neural network

Abstract

fetched live from OpenAlex

The goal of a learning algorithm is to receive a training data set as input and provide a hypothesis that can generalize to all possible data points from a domain set. The hypothesis is chosen from hypothesis classes with potentially different complexity orders that represent controlling parameters in the learning process, also denoted as hyperparameters. Linear regression modeling is an important category of learning algorithms. Uncertainty of target samples in practical applications affects the generalization performance of the learned model. Failing to choose a proper model or hypothesis class can lead to serious issues such as underfitting or overfitting. These issues have been addressed by alternating cost functions or by utilizing cross-validation methods. These approaches can introduce new hyperparameters with their own new challenges and uncertainties or increase the computational complexity of the learning algorithm. On the other hand, the theory of probably approximately correct (PAC) aims at defining learnability based on probabilistic settings. Despite its theoretical value, PAC does not address practical learning issues on many occasions. This thesis is inspired by the foundation of PAC and is motivated by existing regression learning issues. The proposed approach, denoted by ε-Confidence Approximately Correct (ε-CoAC), utilizes Kullback—Leibler divergence (relative entropy) and proposes a new related typical set in the set of hyperparameters to tackle the learnability issue. ε-CoAC learnability is able to validate the learning process as a function of data length and as a function of the complexity order of the hypothesis class. Moreover, it enables the learner to compare hypothesis classes of different complexity order (hyperparameters) and choose among them the optimum with the minimum ε in the ε-CoAC framework. The ε-CoAC learnability not only overcomes the issues of overfitting and underfitting, but also shows advantages and superiority over the well-known cross-validation method in terms of time consumption and in terms of accuracy. A valuable application of ε-CoAC learnability is presented for simultaneous model order and time delay selection for LTI systems. Classical methods have approached this problem from two separate angles for time-delay estimation and for order selection with different cost functions. The ε-CoAC approach solves the problem with a unified cost function. The proposed method not only outperforms existing approaches but is also shown to be more robust to variations of the signal to noise ratio (SNR). The approach is also extended for online impulse response estimation and introduces efficient stopping criteria that are extremely valuable in practical applications. For the second hyperparameter analysis in machine learning, the challenge of regularization hyperparameter selection for the Support Vector Machine (SVM) algorithm is addressed. The regularization parameter controls the model capacity and the trade-off between the training and the generalization errors. It is shown that interestingly the introduced Separability and Scatteredness (S&S) ratio plays a key role in SVM hyperparameter selection, including kernel hyperparameters. Importance of S&S ratio in this context is similar to the role of the signal-to-noise ratio in the signal processing context. The proposed method outperforms existing cross-validation approaches, especially in the sense of computational complexity. For the hyperparameter selection in unsupervised learning, the fundamentals of ε-CoAC learnability is utilized by viewing the problem of clustering from a new angle. The application of the proposed stochastic based hyperparameter selecting algorithm can be generalized in the form of a validity index. The new validity index is shown to be superior to the state-of-the-art validity indices in the sense of accuracy and robustness to the cluster shape. Finally, the proposed validation index approach is extended for application in graph node clustering. The approach shows advantages over the existing methods in the sense of conductance and graph-based normalizing cuts.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.011
metaresearch head score (Gemma)0.045
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Simulation or modeling · Consensus signal: Simulation or modeling
GenreCandidate signal: Empirical · Consensus signal: none
Teacher disagreement score0.011
Threshold uncertainty score0.056

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0110.045
Meta-epidemiology (narrow)0.0010.001
Meta-epidemiology (broad)0.0020.002
Bibliometrics0.0030.002
Science and technology studies0.0010.004
Scholarly communication0.0030.004
Open science0.0020.002
Research integrity0.0020.003
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.017
GPT teacher head0.268
Teacher spread0.251 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designSimulation or modeling
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2024
Admission routes1
Has abstractyes

Explore more

Same topicMachine Learning and Data ClassificationFrench-language works237,207