Bibliographic record
Abstract
Recursion for Adversarial Modelling: New Evidence W. Joseph MacInnes (MacInnes@utsc.utoronto.ca) Centre for Computational and Cognitive Neuroscience, University of Toronto at Scarborough 1265 Military Trail, Scarborough, Ont. M1C 1A4, Canada Debra Gilin (Debra.Gilin@smu.ca) Department of Psychology, Saint Mary's University Halifax, N.S.. B3H 3C3, Canada was instructed to win the most money for their 'country' and was allowed frequent negotiations for strategy. To explore the discrepancy of what we choose to call strategic and personality modelling, both were measured: the personality scale as used in Burns (1998), and a second for strategic modelling (both for recursive levels 0-3). Regression results showed that only personality modelling was significant in predicting how well a subject did in terms of outcome in the game. Further, although depth 2 recursion was primarily responsible for this effect, it was winnings through cooperation which was influenced by this modelling ability. These results replicate Burns (1998), since personality modelling was also used in that study, but would also explain other studies. Computer/Computer matches did show a benefit of strategic modelling since the agents involved were incapable of producing or measuring personality as in the human study. It could also explain the computer/human null result since the computer agent only modelled the human's strategy in the game, and had no model of personality. If deception, however, were the reason for the depth 2 advantage, we would expect money earned from defecting to be influenced, where in fact we see the opposite. Since cooperation (not deception) benefits from depth 2 recursion in this experiment it seems more likely that depth 2 plays a broader role in conflict and negotiation. Introduction The use of recursion in modelling an adversary has been suggested as a crucial component in a number of tasks including game theory, games, negotiations, economics, war and even in the evolution of human intelligence. Thagard (1992) defined recursive modeling (RM) as the ability to place oneself in the mindset of ones opponent, and to do so at different depths. These RM depths consisted of depth 0 - self insight ( I know what I will do the environment ); depth 1 - perspective ( I will include a model of what I believe my opponent will do ); Depth 2 - meta perspective ( I will include what I believe my opponent thinks I will do ); and so on. Thagard suggested that depth 2 held special importance in success against an adversary, since this is where deception would take place. I would need to understand what an adversary thought of my strategy in order or influence or manipulate that belief. A number of studies have looked at RM with a variety of tasks and opponents including many combinations of human and intelligent computer agents. Burns and Vollmeyer (1998) tested human/human dyads in a simple guessing game from game theory and discovered that subjects who were skilled depth two modellers in a questionnaire performed better on the game theory task. MacInnes (2001) incorporated RM in intelligent game agents (computer/computer), and also showed that they could benefit from depth two recursion. The next step (computer agents which could recursively model human behaviour), however met with less success. MacInnes (2004), using a number of intelligent algorithms failed to show a benefit of RM (depth 0 was optimal in most conditions). A number of theories were presented for this result including: a) Machine learning algorithms had already incorporated recursion implicitly from training subjects (presented here, MacInnes 2006). b) Although the theory claimed that RM strategy produced the benefit, it was opponent’s personality modelling which was actually measured in previous human modelling research. Since these theories are not mutually exclusive, a) will be left for future work, and b) will be explored here. Experiment and Results The experiment was a complex game with prisoner's dilemma style payouts. Short term gains could be made through defection, but long term gain could only be achieved through the development of trust. Each participant Acknowledgments Funding provided in part by NSERC Canada, the Centre for Computational and Cognitive Neuroscience (UTSC), and Saint Mary’s University. References Burns, B & Vollmeyer, R (1998). Modelling the Adversary and Success in Competition. Journal of Personality and Social Psychology. (75 No. 3) 711-718. MacInnes, J., Banyasad, O. & Upal, A. (2001). Watching Me, Watching You. Recursive modeling of autonomous agents. Abstracts of the Canadian Conference on AI 2001, Ottawa, Ontario. p 361-364. MacInnes, W.J. (2004) Believability in Multi-Agent Computer Games: Revisiting the Turing Test. Proceedings of CHI, 1537. MacInnes, W.J. (2006). Modelling the enemy: Recursive Cognitive Models in Dynamic Environments. In Press, Cognitive Science Conference. Thagard, P. (1992). Adversarial Problem Solving: Modeling and Opponent Using Explanatory Coherence. Cognitive Science.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.053 | 0.294 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.003 | 0.003 |
| Bibliometrics | 0.003 | 0.003 |
| Science and technology studies | 0.002 | 0.010 |
| Scholarly communication | 0.008 | 0.021 |
| Open science | 0.008 | 0.006 |
| Research integrity | 0.006 | 0.016 |
| Insufficient payload (model declined to judge) | 0.049 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".