Learning verb alternations in a usage-based Bayesian model - eScholarship
Bibliographic record
Abstract
Learning verb alternations in a usage-based Bayesian model Christopher Parisien and Suzanne Stevenson Department of Computer Science, University of Toronto Toronto, ON, Canada {chris, suzanne}@cs.toronto.edu Abstract One of the key debates in language acquisition involves the degree to which children’s early linguistic knowledge employs abstract representations. While usage-based accounts that fo- cus on input-driven learning have gained prominence, it re- mains an open question how such an approach can explain the evidence for children’s apparent use of abstract syntactic gen- eralizations. We develop a novel hierarchical Bayesian model that demonstrates how abstract knowledge can be generalized from usage-based input. We demonstrate the model on the learning of verb alternations, showing that such a usage-based model must allow for the inference of verb class structure, not simply the inference of individual constructions, in order to account for the acquisition of alternations. Keywords: Verb learning; language acquisition; Bayesian modelling; computational modelling. Introduction An important debate in language acquisition concerns the na- ture of children’s early syntax. On one side of the debate lies a claim that children develop their syntactic knowledge in an item-based manner. This claim of usage-based learning ar- gues that very young children associate verb argument struc- ture with specific lexical items, only gradually abstracting syntactic knowledge after four years of age (e.g., Tomasello, 2003). An alternative claim suggests that young children do indeed possess abstract syntactic representations—i.e., gen- eralizations about the structure of their language that are not necessarily tied to lexical items (e.g., Fisher, 2002). Syntactic alternation structure is often considered to be a central phenomenon in this debate. Consider the following example of the English dative alternation: (1) I gave a toy to my dog. (2) I gave my dog a toy. These sentences mean roughly the same thing, but are ex- pressed in different ways. The first, a prepositional dative, expresses the theme (a toy) as an object and the recipient (my dog) in a prepositional phrase. The second, a double-object dative, expresses both the theme and recipient as objects and reverses their order. Verbs that allow similar alternations often have similar se- mantics (Levin, 1993), which suggests that alternations re- flect much of our cognitive representations of verbs. Fur- thermore, these regularities appear to influence our language use. In word learning experiments, children as young as three years of age appear to use abstract representations of the da- tive alternation (Conwell & Demuth, 2007). While this is ev- idence of abstract syntax at a very young age, it does not nec- essarily invalidate the usage-based hypothesis, since the ab- stractions may originate from item-specific representations. One way to bring these opposing positions together is to demonstrate, using naturalistic data, how to connect a usage- based representation of language with abstract syntactic gen- eralizations. We argue that alternation structure can be ac- quired and generalized from usage patterns in the input, with- out a priori expectations of which alternations may or may not be acceptable in the language. We support this claim us- ing a hierarchical Bayesian model (HBM) which is capable of making inferences about verb argument structure at multiple levels of abstraction simultaneously. We show that the in- formation relevant to verb alternations can be acquired from observations of how verbs occur with individual arguments in the input. In this sense, we present a competency model showing what can be acquired, but we do not make claims regarding the specific processing mechanisms involved. From a corpus of child-directed speech, our model acquires a wide variety of argument structure constructions over hun- dreds of verbs. Moreover, by forming classes of verbs with similar usage patterns, the model can generalize knowledge of alternation patterns to novel verbs. This stands in contrast to earlier models which have focused on either the acquisition of the constructions themselves, or the formation of classes over given constructions. The integration in our model of these two important aspects of verb learning has implications for current theories of language acquisition, by showing how abstract syntactic knowledge can be acquired and generalized from usage-level input. Related work Previous computational approaches to language acquisition have used HBMs to represent the abstract structure of verb use. Alishahi and Stevenson (2008) used an incremental Bayesian model to cluster individual verb usages (or tokens), simulating the acquisition of verb argument structure con- structions. Using naturalistic input, the authors showed how a probabilistic representation of constructions can explain chil- dren’s recovery from overgeneralization errors. In another Bayesian model of verb learning, Perfors et al. (2010) clus- ter verb types by comparing the variability of constructions for each of the verbs. The model can distinguish alternating from non-alternating dative verbs and can make appropriate generalizations when learning novel verbs. Both of the above models show realistic patterns of gen- eralization, but they operate at complementary levels of ab- straction. The model of Alishahi and Stevenson does not cap- ture the alternation patterns of verbs, while Perfors et al. as- sume that the individual constructions participating in the al- ternation have already been learned. Furthermore, Perfors et
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.003 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".