Bibliographic record
Abstract
This paper examines the occurrence of epenthetic vowels before and between the initial consonant clusters in Bengali speakers of English, and provides an Optimality Theory (OT) analysis to account for this phenomenon. Native Bengali words disallow initial consonant clusters, and many word-initial consonant clusters in loan words are simplified according to these phonotactics. The maximum syllabic structure is CVC in Bengali and speakers often carry this restriction over to loan words. For example, geram (CV.CVC) instead of gram (CCVC) for the Sanskrit loan word village, or iskul (VC.CVC) instead of skul (CCVC) for the English word school (Kar, 2009). I argue that in rising sonority clusters, a vowel is inserted between the two consonants and in falling sonority clusters (i.e., [s]-stop clusters) the vowel is inserted before the consonant cluster. I also explain that the sonority sequencing constraint SYLLABLE CONTACT treats [s]-stop clusters differently from obstruent-sonorant clusters, and the differing epenthesis pattern can be explained properly if it is considered an effect of SYLLABLE CONTACT – the preference of sonority to fall across a syllable boundary, which was proposed by Murray and Venneman (1983) and also supported by Gouskova (2001). With tableaux, I demonstrate that the epenthesis in consonant clusters is caused by the prohibition on consonant clusters in Bengali and the site of epenthesis is determined by SYLLABLE CONTACT (Gouskova, 2001). I also demonstrate that the constraint that prefers epenthesis before the [s]stop cluster is CONTIG-IO (Kager, 1999). Furthermore, I propose that apart from SYLLABLE CONTACT, two other constraints *OO and *OR can also account for the vowel epenthesis in Bengali.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.002 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".