Recursion for Adversarial Modeling: New Evidence
Notice bibliographique
Résumé
Recursion for Adversarial Modelling: New Evidence W. Joseph MacInnes (MacInnes@utsc.utoronto.ca) Centre for Computational and Cognitive Neuroscience, University of Toronto at Scarborough 1265 Military Trail, Scarborough, Ont. M1C 1A4, Canada Debra Gilin (Debra.Gilin@smu.ca) Department of Psychology, Saint Mary's University Halifax, N.S.. B3H 3C3, Canada was instructed to win the most money for their 'country' and was allowed frequent negotiations for strategy. To explore the discrepancy of what we choose to call strategic and personality modelling, both were measured: the personality scale as used in Burns (1998), and a second for strategic modelling (both for recursive levels 0-3). Regression results showed that only personality modelling was significant in predicting how well a subject did in terms of outcome in the game. Further, although depth 2 recursion was primarily responsible for this effect, it was winnings through cooperation which was influenced by this modelling ability. These results replicate Burns (1998), since personality modelling was also used in that study, but would also explain other studies. Computer/Computer matches did show a benefit of strategic modelling since the agents involved were incapable of producing or measuring personality as in the human study. It could also explain the computer/human null result since the computer agent only modelled the human's strategy in the game, and had no model of personality. If deception, however, were the reason for the depth 2 advantage, we would expect money earned from defecting to be influenced, where in fact we see the opposite. Since cooperation (not deception) benefits from depth 2 recursion in this experiment it seems more likely that depth 2 plays a broader role in conflict and negotiation. Introduction The use of recursion in modelling an adversary has been suggested as a crucial component in a number of tasks including game theory, games, negotiations, economics, war and even in the evolution of human intelligence. Thagard (1992) defined recursive modeling (RM) as the ability to place oneself in the mindset of ones opponent, and to do so at different depths. These RM depths consisted of depth 0 - self insight ( I know what I will do the environment ); depth 1 - perspective ( I will include a model of what I believe my opponent will do ); Depth 2 - meta perspective ( I will include what I believe my opponent thinks I will do ); and so on. Thagard suggested that depth 2 held special importance in success against an adversary, since this is where deception would take place. I would need to understand what an adversary thought of my strategy in order or influence or manipulate that belief. A number of studies have looked at RM with a variety of tasks and opponents including many combinations of human and intelligent computer agents. Burns and Vollmeyer (1998) tested human/human dyads in a simple guessing game from game theory and discovered that subjects who were skilled depth two modellers in a questionnaire performed better on the game theory task. MacInnes (2001) incorporated RM in intelligent game agents (computer/computer), and also showed that they could benefit from depth two recursion. The next step (computer agents which could recursively model human behaviour), however met with less success. MacInnes (2004), using a number of intelligent algorithms failed to show a benefit of RM (depth 0 was optimal in most conditions). A number of theories were presented for this result including: a) Machine learning algorithms had already incorporated recursion implicitly from training subjects (presented here, MacInnes 2006). b) Although the theory claimed that RM strategy produced the benefit, it was opponent’s personality modelling which was actually measured in previous human modelling research. Since these theories are not mutually exclusive, a) will be left for future work, and b) will be explored here. Experiment and Results The experiment was a complex game with prisoner's dilemma style payouts. Short term gains could be made through defection, but long term gain could only be achieved through the development of trust. Each participant Acknowledgments Funding provided in part by NSERC Canada, the Centre for Computational and Cognitive Neuroscience (UTSC), and Saint Mary’s University. References Burns, B & Vollmeyer, R (1998). Modelling the Adversary and Success in Competition. Journal of Personality and Social Psychology. (75 No. 3) 711-718. MacInnes, J., Banyasad, O. & Upal, A. (2001). Watching Me, Watching You. Recursive modeling of autonomous agents. Abstracts of the Canadian Conference on AI 2001, Ottawa, Ontario. p 361-364. MacInnes, W.J. (2004) Believability in Multi-Agent Computer Games: Revisiting the Turing Test. Proceedings of CHI, 1537. MacInnes, W.J. (2006). Modelling the enemy: Recursive Cognitive Models in Dynamic Environments. In Press, Cognitive Science Conference. Thagard, P. (1992). Adversarial Problem Solving: Modeling and Opponent Using Explanatory Coherence. Cognitive Science.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,053 | 0,294 |
| Méta-épidémiologie (sens strict) | 0,002 | 0,002 |
| Méta-épidémiologie (sens large) | 0,003 | 0,003 |
| Bibliométrie | 0,003 | 0,003 |
| Études des sciences et des technologies | 0,002 | 0,010 |
| Communication savante | 0,008 | 0,021 |
| Science ouverte | 0,008 | 0,006 |
| Intégrité de la recherche | 0,006 | 0,016 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,049 | 0,006 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».