ROBUSTNESS INSTEAD OF ACCURACY SHOULD BE THE PRIMARY OBJECTIVE FOR SUBJECTIVE PATTERN RECOGNITION RESEARCH: STABILITY ANALYSIS ON MULTICANDIDATE ELECTORAL COLLEGE VERSUS DIRECT POPULAR VOTE
Notice bibliographique
Résumé
Subjective pattern recognition is a class of pattern recognition problems, where we not only merely know a few, if any, the strategies our brains employ in making decisions in daily life but also have only limited ideas on the standards our brains use in determining the equality/inequality among the objects. Face recognition is a typical example of such problems. For solving a subjective pattern recognition problem by machinery, application accuracy is the standard performance metric for evaluating algorithms. However, we indeed do not know the connection between algorithm design and application accuracy in subjective pattern recognition. Consequently, the research in this area follows a “trial and error” process in a general sense: try different parameters of an algorithm, try different algorithms, and try different algorithms with different parameters. This phenomenon can be observed clearly in the nearly 30 years research of the face recognition: although huge advances have been made, no algorithm has ever been shown a potential to be consistently better than most of the algorithms developed earlier; it was even shown that a naïve algorithm can work, in the sense of accuracy, at least no worse than many newly developed ones in a few benchmarks. We argue that, the primary objective of subjective pattern recognition research should be moved to theoretical robustness from application accuracy so that we can evaluate and compare algorithms without or with only few “trial and error” steps. We in this paper introduce an analytical model for studying the theoretical stabilities of multicandidate Electoral College and Direct Popular Vote schemes (aka regional voting scheme and national voting scheme, respectively), which can be expressed as the a posteriori probability that a winning candidate will continue to be chosen after the system is subjected to noise. This model shows that, in the context of multicandidate elections, generally, Electoral College is more stable than Direct Popular Vote, that the stability of Electoral College increases from that of Direct Popular Vote as the size of the subdivided regions decreases from the original nation size, up to a certain level, and then the stability starts to decrease approaching the stability of Direct Popular Vote as the region size approaches the original unit cell size; and that the stability of Electoral College approaches that of Direct Popular Vote in the two extremities as the region size increases to the original national size or decreases to the unit cell size. It also shows a special situation of white noise dominance with negligibly small concentrated noise, where Direct Popular Vote is surprisingly more stable than Electoral College, although the existence of such a special situation is questionable. We observe that “high stability” in theory indeed always reveals itself in “high accuracy” in applications. Extensive experiments on two human face benchmark databases applying an Electoral College framework embedded with standard baseline and newly developed holistic algorithms have been conducted. The impressive improvement by Electoral College over regular holistic algorithms verifies the stability theory on the voting systems. It also shows an evidential support for adopting theoretical stability instead of application accuracy as the primary objective for subjective pattern recognition research.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,011 | 0,094 |
| Méta-épidémiologie (sens strict) | 0,001 | 0,000 |
| Méta-épidémiologie (sens large) | 0,001 | 0,001 |
| Bibliométrie | 0,002 | 0,002 |
| Études des sciences et des technologies | 0,001 | 0,005 |
| Communication savante | 0,003 | 0,008 |
| Science ouverte | 0,002 | 0,003 |
| Intégrité de la recherche | 0,002 | 0,003 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,005 | 0,001 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».