Do Dogs Prefer Helpers in an Infant-Based Social Evaluation Task?
Bibliographic record
Abstract
Social evaluative abilities emerge in human infancy, highlighting their importance in shaping our species' early understanding of the social world. Remarkably, infants show social evaluation in relatively abstract contexts: for instance, preferring a wooden shape that helps another shape in a puppet show over a shape that hinders another character (Hamlin et al., 2007). Here we ask whether these abstract social evaluative abilities are shared with other species. Domestic dogs provide an ideal animal species in which to address this question because this species cooperates extensively with conspecifics and humans and may thus benefit from a more general ability to socially evaluate prospective partners. We tested dogs on a social evaluation puppet show task originally used with human infants. Subjects watched a helpful shape aid an agent in achieving its goal and a hinderer shape prevent an agent from achieving its goal. We examined (1) whether dogs showed a preference for the helpful or hinderer shape, (2) whether dogs exhibited longer exploration of the helpful or hinderer shape, and (3) whether dogs were more likely to engage with their handlers during the helper or hinderer events. In contrast to human infants, dogs showed no preference for either the helper or the hinderer, nor were they more likely to engage with their handlers during helper or hinderer events. Dogs did spend more time exploring the hindering shape, perhaps indicating that they were puzzled by the agent's unhelpful behavior. However, this preference was moderated by a preference for one of the two shapes, regardless of role. These findings suggest that, relative to infants, dogs show weak or absent social evaluative abilities when presented with abstract events and point to constraints on dogs' abilities to evaluate others' behavior.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".