Statistical learning in children's emergent L2 literacy: Cross-cultural insights from rural Côte d'Ivoire
Bibliographic record
Abstract
Studies of non-linguistic statistical learning (SL) have often linked performance in SL tasks with differences in language outcomes. Most of these studies have focused on Western and high-income educational contexts, but children worldwide learn in radically different educational systems and communities, and often in a second language. In the west African nation of Côte d’Ivoire, children enter fifth grade (CM-1) with widely varying ages and literacy skills. Across three iteratively-developed experiments, 157 children, age 8-15 years, in rural communities in the greater-Adzópe region of Côte d’Ivoire watched sequences of cartoon images with embedded triplet patterns on touchscreen tablets, while performing a target-detection task. We assessed these tablet-based adaptations of non-linguistic visual SL and asked whether the children’s individual differences in performance on the SL tasks were related to their first and second language and literacy skills. We found group-level evidence that children used the statistical regularities in the image sequence to gradually decrease their response times, but their responses on post-test discrimination did not reflect this learning. When evaluating the correlation between SL and language skills, individual differences related to other task demands predicted oral language skills shared by first and second languages, while SL better predicted second language print skills. These findings suggest that non-linguistic SL paradigms can measure similar skills in Ivorian children as previous samples, but they also echo recent calls for further cross-cultural validation, greater internal reliability, and tests for confounding variables (such as processing speed) in studies of individual differences in statistical learning.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.004 |
| Meta-epidemiology (narrow) | 0.001 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.002 | 0.001 |
| Science and technology studies | 0.002 | 0.002 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.001 | 0.002 |
| Research integrity | 0.001 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".