Categorical and non-categorical perception of marginal phonemes
Notice bibliographique
Résumé
Marginal phonemes and contrasts occupy a complex position in linguistic theory, as traditional theories of phonemehood do not account for marginality. However, contemporary linguistics has found that phonemic contrast strength is not fixed in childhood but rather continues to change, thus implicating the lexicon, the set of words a speaker knows, as a factor in the behavior of phonemic contrasts.This dissertation takes this link between category strength and the lexicon and treats it as an empirical question. I identify token frequency and type informativity, measures of frequency and predictability within a lexicon, as potential predictors of individual behavior. I then justify and present an experimental procedure for an eye tracking, two-alternative forced choice, categorization study on three phonetic continua — [a͡ɪ]-[ʌ͡i], a marginal contrast; [a͡ɪ]-[ɔ͡ɪ], a classic phonemic contrast; and [ʌ͡i]-[ɔ͡ɪ], a mixed case — in Canadian English, using the visual world paradigm. I discuss decisions that were made in the design of the experiment, including how individual lexicons were probed and why multiple continua were examined.\nAnalyzing the resultant eye tracking data both graphically and by GAMM model comparison, I find that behavior was not interpretably predicted by my selected predictors, though their contribution to bias, a normalized preference measure, was statistically significant. I report on the behavioral patterns that were found in the categorization data and show that participants with differential behavior did not have statistically significant differences in either frequency or informativity in nearly all cases.\nMy findings come as a surprise, as predicting that variation in the lexicon (operationalized as frequency and informativity) should influence linguistic behavior is both obvious and supported by the literature. I thus present my thoughts on why these predictors were not significant ones as well as my suspicion that the process of calculating these lexical statistics was poisoned by the likely incorrect assumption that a marginal phoneme can be treated as if it were a strong phoneme for the purposes of calculation. I close with suggestions for future work that could advance understanding of this issue, including potential test cases and the need for alternative operationalizations of frequency and predictability for marginal phonemes.
Récupéré en direct depuis OpenAlex et désinversé. Les résumés ne sont pas conservés dans cette base de données : les index inversés représentent 8,6 Go des 9,3 Go de texte de la base, et le serveur dispose de 13 Go libres.
Comment cette classification a été obtenuedéplier
Prédiction machine sur la base complète
Imitation des enseignantsNi prévalence calibrée, ni vérité terrain. Validation humaine à venir. Le volet Gemma est une étiquette directe du modèle pour chaque travail de la base, lue sur la notice réduite au titre. Le volet Codex est un classifieur appris des 10 348 étiquettes directes de Codex et calibré sur les taux pondérés de l'échantillon; les champs sans appui suffisant ne portent aucun appel Codex. Le mode candidate est l'union des deux volets; le consensus est leur intersection. Ces sorties portent le statut machine_predicted_unvalidated et ne sont pas des étiquettes humaines.
Scores du classifieur distillé par catégorie (deux têtes)
| Catégorie | Codex | Gemma |
|---|---|---|
| Métarecherche | 0,001 | 0,007 |
| Méta-épidémiologie (sens strict) | 0,000 | 0,000 |
| Méta-épidémiologie (sens large) | 0,000 | 0,000 |
| Bibliométrie | 0,001 | 0,000 |
| Études des sciences et des technologies | 0,000 | 0,001 |
| Communication savante | 0,002 | 0,001 |
| Science ouverte | 0,000 | 0,001 |
| Intégrité de la recherche | 0,001 | 0,001 |
| Charge utile insuffisante (le modèle a refusé de juger) | 0,003 | 0,000 |
Scores machine (provisoires)
Les deux têtes enseignantes du modèle étudiant, lues sur ce travail. Un score ordonne la base pour la relecture; il n'affirme jamais une catégorie, et le statut de validation accompagne chaque rangée tel quel.
Scores de référence d'un modèle non mature (critères de maturité non atteints, 7 itérations). Un score ordonne; il n'affirme jamais une catégorie.
score_only:v0-immature-baseline · tel quel depuis la passe de notation : score_only signifie que le nombre peut ordonner les travaux, et qu'aucune étiquette de catégorie n'en découleClassification
machine, non validéePrédiction automatique; un appel candidat d’une seule source (Gemma direct ou Codex distillé), pas un consensus.
Le détail, modèle par modèle et score par score, se trouve en fin de page sous « Comment cette classification a été obtenue ».