The more they hear the more they learn? Using data from bilinguals to test models of early lexical development
Bibliographic record
Abstract
Children have an early ability to learn and comprehend words, a skill that develops as they age. A critical question remains regarding what drives this development. Maturation-based theories emphasise cognitive maturity as a driver of comprehension, while accumulator theories emphasise children’s accumulation of language experience over time. In this study we used archival looking-while-listening data from 155 children aged 14–48 months with a range of exposure to the target languages (from 10% to 100%) to evaluate the relative contributions of maturation and experience. We compared four statistical models of noun learning: maturation-only, experience-only, additive (maturation plus experience), and accumulator (maturation times experience). The best-fitting model was the additive model in which both maturation (age) and experience were independent contributors to noun comprehension: older children as well as children who had more experience with the target language were more accurate and looked faster to the target in the looking-while-listening task. A 25% change in relative language exposure was equivalent to a 4 month change in age, and age effects were stronger at younger than at older ages. Whereas accumulator models predict that the lexical development of children with less exposure to a language (as is typical in bilinguals) should fall further and further behind children with more exposure to a language (such as monolinguals), our results indicate that bilinguals are buffered against effects of reduced exposure in each language. This study shows that continuous-level measures from individual children’s looking-while-listening data, gathered from children with a range of language experience, provide a powerful window into lexical development.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.000 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.004 | 0.007 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".