Fuzzy Classification Using Pattern Discovery
Notice bibliographique
Résumé
Rule-based classifiers allow rationalization of classifications made. This in turn improves understanding which is essential for effective decision support. As a rule based classifier, the pattern discovery (PD) algorithm functions well in discrete, nominal and continuous data domains. A drawback when using PD as a classifier for decision support is that it has an unbounded decision space that confounds the understanding of the degree of support for a decision. Incorporating PD into a fuzzy inference system (FIS) allows the the degree of support for a decision to be expressed with intuitively understandable terms. In addition, using discrete algorithms in continuous domains can result in reduced accuracy due to quantization. Fuzzification reduces this ldquocost of quantizationrdquo and improves classification performance. In this work, the PD algorithm was used as a source of rules for a series of FISs implemented using different rule weighting and defuzzification schemes, each providing a linguistic basis for rule description and a bounded space for expression of decision support. The output of each FIS consists of a suggested outcome, a strong confidence metric describing suggestions within this space and a linguistic expression of the rules. This constitutes a stronger basis for decision making than that provided by PD alone. A variety of synthetic, continuous class distributions with varying degrees of separation was used to evaluate the performance of fuzzy, PD, back-propagation and Bayesian classifiers. Overall, the accuracy of the fuzzy system was found to be similar, but slightly below, that of the inherently continuous valued classifiers and was somewhat improved with respect to the PD classifiers. For the difficult spiral class distributions studied, the fuzzy classifiers were able to make more classifications than the PD classifiers. The correct classification rates for the fuzzy classifiers were similar across the various rule weighting and defuzzification schemes, demonstrating the strength of the statistical method for rule generation. Analysis of several real-world data sets shows that a PD-based FIS has comparable performance to a neuro-fuzzy system. The use of a PD based FIS however, provides insight into the structure of the data analyzed not available through the other approaches.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,004 | 0,010 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,001 |
| Méta-épidémiologie (sens large) | 0,002 | 0,002 |
| Bibliométrie | 0,006 | 0,005 |
| Études des sciences et des technologies | 0,001 | 0,001 |
| Communication savante | 0,005 | 0,003 |
| Science ouverte | 0,002 | 0,002 |
| Intégrité de la recherche | 0,002 | 0,002 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,005 | 0,003 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».