Phonological redeployment and the mapping problem: Cross-linguistic E-similarity is the beginning of the story, not the end
Bibliographic record
Abstract
In this research note I want to address some misunderstandings about the construct of redeployment and suggest that we need to fit these behavioural data from Yang, Chen and Xiao (YCX) into a broader context. I will suggest that these authors’ work is not just about the failure of three models to predict equivalence classification. Equivalence classification is not the end of the story but only the beginning. We need to look at what cues are detected in the input, which subset of the input becomes intake, and how this intake is parsed onto phonological structures. The empirical results of YCX should not be viewed as some sort of non-result inasmuch as none of the proposed predictors of Mandarin equivalence classification foresaw that the Russian prevoiced stops and short-lag stops would be equated with the Mandarin short-lag stops. Rather, the empirical results need to be contextualized by considering such factors as cue reweighting as part of the learning theory which maps intake onto phonological representations. In this light, the results are not a repudiation of phonological redeployment, but help to shed light on the parsing of the acoustic signal, the importance of robust burst-release cues, and the non-local nature of L2 phonological learning (as opposed to noticing).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.055 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.006 |
| Scholarly communication | 0.004 | 0.019 |
| Open science | 0.002 | 0.004 |
| Research integrity | 0.002 | 0.005 |
| Insufficient payload (model declined to judge) | 0.009 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".