The nature of regressions in the acquisition of phonological grammars
Bibliographic record
Abstract
Children’s acquisition of their L1 phonological grammar is typically understood as a gradual progression from an initial universal state towards a language-specific one, in which learners incrementally change their grammars to better approximate the target. One challenging problem for this view, however, are the many reports of ‘U-shaped development’ in which production temporarily regresses, diverging further from the target rather than drawing closer. Based on existing and novel analyses of longitudinal data, this paper argues that phonological regressions should not be captured directly within the normal workings of children’s error-driven mechanisms for grammar learning. It also identifies a kind of regression that seems plausible but is nonetheless apparently unattested: one in which markedness constraints flip-flop over time, so that improvement on one marked structure entails regression on another. With this initial empirical base, the paper then demonstrates that an error-driven OT-like learner which stores its errors and imposes certain persistent biases can in fact easily regress in the unattested way. Section 5 discusses how OT’s grammatical parallelism is in part responsible for creating the unattested regression pattern, and how a serial constraint-based grammar like Harmonic Serialism (McCarthy 2007 et seq) avoids this regression.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.011 |
| Meta-epidemiology (narrow) | 0.000 | 0.001 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.000 |
| Science and technology studies | 0.000 | 0.003 |
| Scholarly communication | 0.002 | 0.003 |
| Open science | 0.001 | 0.001 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.002 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".