Morphological Variation in Haitian Creole
Bibliographic record
Abstract
Among French-based creole languages, Haitian Creole is the one with the highest degree of standardization. The written norm, Standard Haitian Creole (SHC), is based on the speech of monolinguals of the capital area, Port-au-Prince, rather than on the variety (kreyòl swa) of the politically and economically powerful Creole–French bilingual minority. For instance, the front rounded vowels and postvocalic /r/ of the latter are absent from SHC, which is spreading to the rest of the country through the media and the educational system. \n \nIn order to evaluate the diffusion of SHC, a sociolinguistic study of Northern Haitian Creole (Capois) was conducted in and around Cape Haitian, whose spoken variety diverges most from SHC. In addition to stereotypical features such as the possessive kin a + pronoun (vs. SHC pa + pronoun), we uncovered several Capois features—some of which were first described in Étienne (1974)—still in widespread use in Northern Haiti. In this article, we focus on the most frequently occurring variable, the third person singular pronoun (3SG), which alternates between SHC li/l, and Capois i/y. \n \nUsing a corpus of 24 speakers, we show that SHC li has yet to replace Capois i, which is preferred by a large proportion of community members (90%; N=2,823) and used categorically in the existential context i gen ‘there is/are’. For both the rural and urban populations, this variable is conditioned by syntactic and phonological factors. In subject position, the Capois or the full SHC variants are favored before a consonant, while the reduced SHC form l is the only variant favored before a vowel. In object position, Capois and SHC variants are in near perfect complementary distribution: the Capois variant occurs (near-)categorically after a vowel, and SHC variants occur (near-)categorically after a consonant. Despite these shared tendencies, we found a lower rate of Capois variant use in urban speakers, which may be due to their greater exposure to speakers from other areas of Haiti, to the media (especially television), and closer contact with middle class bilingual speakers who are more influencedby the standard emanating from Port-au-Prince.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.002 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.004 | 0.001 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".