Playing Games with Robots - A Method for Evaluating Human-Robot Interaction
Bibliographic record
Abstract
From the description of our game-based HRI testbed above, we would like to provide some lessons learned in terms of the benefits and challenges of our approach and application. First and foremost, this testbed is relatively simple to construct and cost effective. Using readily available products such as the AIBO and the RolaBoardTM, we were able to rapidly construct and prototype our testbed. This again speaks to the flexible nature of games which can be created with whatever is easily assessable or modified to make implementation easier. Second, because the game has simple and well defined rules and is played within a bounded environment, we can rapidly prototype new games and design new user studies. In fact, we are currently in the second iteration of our testbed which will feature a slightly modified game used to investigate a different research question. Third, we found the use of both physical and virtual entities to be useful for experimental design. Humans currently still have the advantage when it comes to interaction in the physical world. Robots, however, have the advantage when it comes to interacting with digital information. By playing with these factors, various social relationships can be generated such as trust. Finally, although our initial goal with the current testbed was to look at collaboration, we found that having a game which involves collaboration is beneficial for increasing the potential of other forms of social interaction. For example, when humans and robots need to collaborate, they are required to communicate with each other in a much more complex manner than simple command and execution. Certainly, there are a couple of stumbling blocks with our testbed exploration as well. The biggest problem with using games for experimentation is that we can not really control how the game is played. Each participant will have a different gameplay experience based on the outcome of the game and the way the game was played. Therefore, it is difficult to compare data. For example, on the issue of trust, the outcome of the game played significantly affects the participant's opinion since winning tends to build trust. Scripting games is one solution to this problem, but this leads to the dilemma of having to disguise the scripting process to the participant, and in some situations, scripting is not possible. The other problem with our game-based testbed is that evaluation of game experiences in general is a difficult problem by itself. Generally it is hard to collect quantitative data for games, and most forms of gameplay evaluations are often vague. With exploratory studies, these issues are not critical, but with more focused studies, they can skew the data. Currently, we can not offer great solutions to these problems, but we are looking at methods to strike a balance between restrictive and more freeform styles of games. We also have not attempted to transfer the primitive results from our user study to other applications since more rigorous experimentation needs to be performed, but we feel the few results that we do have make sense for other applications as well. However, it is promising to see that the game-based testbed approach is able to explore critical social issues of human-robot interaction such as trust which can assist robot designers in developing future domestic and sociable robots.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.010 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.006 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".