The Influence of Kinetographic Gestures on Real-Time Referential Interpretation
Bibliographic record
Abstract
Considerable evidence suggests listeners can draw on information conveyed by co-speech gestures, even if speakers do not consciously intend their gestures to be communicative. However, the uptake of gesture cues must evaluated on a case-by-case basis, with attention to the specific type of gesture and the kind of information being conveyed. An additional consideration is the extent to which gestures convey information that is distinct from or with speech information.We explore how gestures (reflecting the hand postures required when picking up objects) influence the identification of referents in the immediate context. Apart from examining the semantic effects of this gesture type, the study tests whether co-speech gesture that is in some global sense may nonetheless facilitate comprehension because of how gestures are temporally aligned with the unfolding speech stream. E.g., kinetographs produced in synchrony with verbs could plausibly narrow referential possibilities for upcoming noun phrases by limiting referents to those whose properties match the action reflected in the gesture. This would streamline interpretive processes even when information in the subsequent noun phrase would clearly be sufficient for referent identification.Listeners watched videoclips in which a speaker provided instructions to pick up a scene object (e.g., Pick up the glass), located among other objects with varying size/shape features. On some trials, the verb was accompanied by a kinetographic gesture reflecting the size/shape of the object (in principle enabling anticipation of the correct referent before the noun phrase).Experiment 1 used a gating methodology in which successively longer samples of each videoclip were presented. The first sample ended just before noun onset, and each successive sample was played for 200ms longer than the previous one. After each sample, participants were asked to judge the identity of the referent object. Results showed the presence of gestures conferred a consistent advantage on referent identification, allowing the correct referent to be identified significantly earlier than when gestures were absent. However, one concern with this technique is ecological validity: The task encourages participants to closely attend to the audiovisual stimulus, and the consistent repetition of interrupted speech could create artifactual results.In Experiment 2, videoclips were presented in an uninterrupted format and the task did not require explicit judgments. Participants' eye fixations to scene objects were recorded as the videoclips played in real time. Results showed that gesture continued to confer an advantage on referent identification, although this advantage was comparatively brief within the full timescale of processing. Further, facilitation occurred primarily when gesture information allowed a large object to be differentiated from smaller objects, and not the reverse. Separate analyses explored the extent to which gestures are perceived via direct fixation vs. parafoveally.Overall, the results provide detailed insights into when and how co-speech gesture influences real-time referential interpretation, and also provide a concrete demonstration of how technically redundant gestures can nonetheless facilitate comprehension as sentences unfold.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.018 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.002 | 0.002 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.005 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".