Learning a multi-criteria classification method using machine learning & metaheuristics techniques
Notice bibliographique
Résumé
Data classification is a widely used approach in the area of data mining. Methodologies for addressing data classification have been developed in a variety of research disciplines, including artificial intelligence (AI) and Multi-Criteria Decision Aid (MCDA). The objective of this thesis is to develop a new framework for learning the MCDA method PROAFTN. The limitations of PROAFTN are largely due to the set of parameters required to be obtained to perform the classification procedure. That is, to apply PROAFTN, the values of several parameters need to be determined prior to classification, such as boundaries of intervals and weights. In an MCDA context, these parameters are usually dependent on the judgment of the decision maker (DM). This approach has shortcomings, such as being time consuming and dependent on the availability of a qualified DM. To overcome these limitations and to obtain the best parameters from data, an automatic approach is proposed in this work . This thesis introduces new methodologies based on using machine learning and metaheuristic techniques for establishing PROAFTN parameters from data during the training process. The goal is to obtain from training data the best PROAFTN parameters that achieve the highest classification accuracy. To achieve this, different learning methodologies are proposed in this thesis. Firstly, discretization techniques and an inductive approach are introduced to obtain the required parameters for PROAFTN. Secondly, a different approach based on metaheuristic/hybrid-metaheuristic algorithms is used to develop PROAFTN parameters. The use of metaheuristics to learn PROAFTN begins with the formulation of the optimization problem. Then, population-based methods, namely Particle Swarm Optimization (PSO) and Differential Evolution (DE), and the single-point search method Reduced Variable Neighborhood Search (RUNS) are utilized to obtain the best PROAFTN parameters that can be applied on unseen datasets. To test the performance of the proposed learning approaches, their effectiveness in classification is evaluated on several public-domain datasets and compared to a number of well-known machine learning classifiers. Advanced statistical tests such as the Friedman and Nemenyi tests are used for more meaningful comparisons. The general comparative study, including computational results, demonstrates that the proposed approaches are very competitive with and outrank widely used classification algorithms.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,004 | 0,005 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,002 | 0,002 |
| Bibliométrie | 0,002 | 0,002 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,002 | 0,001 |
| Science ouverte | 0,002 | 0,001 |
| Intégrité de la recherche | 0,002 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,003 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».