Success in Theory of Mind - eScholarship
Bibliographic record
Abstract
Success in Theory of Mind Rose M. Scott (rmscott2@uiuc.edu) Department of Psychology, University of Illinois at Urbana-Champaign, 603 East Daniel St., Champaign, IL 61820 USA Adam R. Petrashek* (arpetras@uwaterloo.ca) Department of Psychology, University of Waterloo, 200 University Ave, Waterloo, ON N2L 3G1 Canada Noah D. Goodman (ndg@mit.edu) Department of Brain and Cognitive Sciences, Massachusetts Institute of Technology, 77 Massachusetts Ave, Cambridge, MA 02139 USA Rebecca Saxe** (saxe@mit.edu) Department of Brain and Cognitive Sciences, Massachusetts Institute of Technology, 77 Massachusetts Ave, Cambridge, MA 02139 USA * denotes organizer, ** denotes discussant Recent evidence suggests that infants in the second year of life can represent a variety of different false beliefs, as well as reason about false perceptions and deception (e.g., Baillargeon, Scott, & He, in press). If infants can represent false beliefs, then why do children fail standard tasks until age 4? Here we argue that this discrepancy reflects the use of different responses. Traditional tasks require children to answer a direct question about an agent's false belief (elicited-response tasks), whereas recent tasks measure children's spontaneous reactions to a scene (spontaneous- response tasks). Simultaneously representing a false belief and planning a response may be too difficult for young children. Since spontaneous tasks do not require a planned response, children succeed much earlier. To examine this possibility, we tested 2.5-year-olds in a novel false-belief task that closely matched the demands of standard tasks but did not require answering a question. While viewing a picture book, children heard a story about an agent who hid her apple in one of two locations; in her absence, the apple was moved to the other location. In the test trial, one picture showed the agent searching for her apple where she had originally hidden it, and one picture showed the agent searching for her apple in its current location. Children looked reliably longer at the original- than at the current- location picture, suggesting that they successfully represented the agent’s false belief. We next tested whether 2.5-year-olds could succeed in an elicited-response task if the response component were made easier for them. Specifically, we provided children with practice with the required response (pointing to one of two locations). In each trial, an experimenter either recited a line of the story (story trials) or asked a question (question trials). On story trials, one picture was shown; on question trials, two pictures were shown and the question required the children to point to one of them. In the final trial, children were asked to point to where the agent would look for her apple. Most children pointed to the correct location (e.g., where the agent falsely believed her apple was Keywords: Theory of mind; cognitive development; social cognition; executive function; social learning; domain- specificity; probabilistic modeling; reaction time. Introduction Peter wants to get the beer that he left in the refrigerator. Predicting Peter’s behaviour correctly is usually an easy matter, but understanding how people correctly predict his behaviour with ease is a much more difficult task. Thirty years of research on theory of mind has focused on the interesting few cases in which people fail to reason about mental states correctly, however it is perhaps more interesting to explore the common, reliable cases of successful theory of mind reasoning. This symposium presents research exploring successful instances of theory of mind reasoning using a variety of experimental approaches, and examines the ability to succeed consistently across the lifespan, with results from toddlers, preschoolers, young children, and adults. Important conclusions are drawn from the presented research, which includes the first evidence that children as young as 2.5 years of age can succeed on explicit false belief tasks (Scott & Baillargeon), the most direct behavioral evidence to date for inhibitory processing in successful behavior prediction based on false belief and avoidance desire in preschoolers and young children (Petrashek & Friedman), and, in adults, evidence from a probabilistic modeling approach to theory of mind and social learning development with extensions to pragmatic language usage and natural pedagogy (Goodman). Why do infants succeed in false-belief tasks when toddlers fail? Evidence for a response account Rose M. Scott & Renee Baillargeon
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.005 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.002 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".