Infant volubility and bilingual input in naturalistic day-long recordings
Bibliographic record
Abstract
Language input plays an important role in child language acquisition. Children encounter language input in different social contexts: They sometimes are directly addressed by caregivers and sometimes overhear conversations between others. For children growing up in bilingual environments, they receive input in two languages. In many cases, a bilingual child is exposed to one language more often than the other which are usually referred to as the child’s dominant and non-dominant languages. Bilingual caregivers are also commonly observed mixing languages when talking to their child. How language input that varies in both quantity and quality interacts with child language development is not fully understood. As well, various units and sampling methods that have been used to measure input can impact the relation between input and language development. To assess bilingually-raised infants’ speech and language development in their first year of life is challenging; in this dissertation, I deployed infant volubility which is a measurement of infants’ vocal activeness and an established precursor of their future language skills. To examine the relation between infant volubility and language input received in different social and language contexts, I analyzed infant and caregiver vocalizations in the naturalistic day-long recordings from the Montréal Bilingual Infants corpus. Twenty-one French-English bilingual families living in Montréal completed language background interviews and contributed three day-long audio recordings when their infant was 10 months old and another day-long audio recording when their infant was 18 months old. Recordings were obtained and processed by the LENA system (Language ENvironment Analysis, Boulder, Colorado). Every other 30-second segment containing adult speech was manually coded for social (one-on-one, overhearing) and language (French, English, mixed-language) contexts. Infant volubility was assessed at two levels (1) overall volubility, the number of infant vocalizations produced throughout the day; and (2) local volubility, the number of infant vocalizations produced in the presence of a certain type of input. In Chapter 2 to 4, I examined infant overall and local volubility’s relation with input in different social contexts (Chapter 2), with English- and French-only input (Chapter 3), and with mixed-language input (Chapter 4). In Chapter 5, I assessed the alignment between input measures estimated by various units (segment counts, adult word counts, speech duration) and different sampling methods (every-other-segment, top-segment).These analyses revealed at least following findings: (1) Infant volubility and language input is robustly and positively related at global and local levels; (2) Infants’ concurrent and longitudinal overall volubility has a strong association with the amount of input received in 1:1 social contexts and in their dominant language; (3) Input received in overhearing contexts and in their non-dominant language also makes unique contributions to infant overall volubility; (4) There is a complex relation between infant volubility and mixed-language input including that a higher proportion of mixed input in 1:1 social contexts is related to reduced overall volubility; (5) Input measures and their relation with infant overall volubility are consistent across different units but diverge across sampling methods.This dissertation adds to an emerging body of research on children’s bilingual language environment and their language acquisition. Findings from this dissertation have important theoretical and practical implications. First, they show a robust relation between infant volubility and language input. Second, they contribute to the current debates on the role of overheard input and mixed-language input. Last, they provide methodological suggestions to future research on the choice of input measurement units and sampling methods
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.004 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".