AM/P~OM/P Merger in Hong Kong vs. Toronto Cantonese: An Under-documented Homeland Sound Change in a Heritage Language Context
Bibliographic record
Abstract
Hong Kong Cantonese has been described as having developed a dissimilatory merger in which O becomes A in pre-labial contexts (henceforth OM/P vs. AM/P). This change is also one that has been described as completed by the end of the 20th century. This paper presents what may be the first acoustic study addressing the Cantonese AM/P~OM/P merger. It also addresses the extent to which heritage speakers in Toronto, Canada also participate in this change. Analysis involved midpoint F1, F2, and F3 measurements from a total of 38 sociolinguistic interviews from the Heritage Language Variation and Change (HLVC) in Toronto Corpus (Nagy 2011). This amounted to a grand total of 889 tokens of AM/P and 816 tokens of OM/P. The Year of Birth of participants ranged from 1922 to 1998. Mixed effects modeling showed that OM/P is significantly raised (lower F1, p < 0.001), significantly retracted (lower F2, p < 0.001), and significantly more rounded (lower F3, p < 0.01) than AM/P. Pillai Scores for each individual speaker were also calculated. A Pearson Correlation test showed a significant inverse correlation between Year of Birth and Pillai Score (r(36) = -0.361, p< 0.05). While no significant difference was found in F1, F2, or F3 variation based on City or based on Generational Group, the overall results from this study show that a merger that was previously described as complete is, in fact, still ongoing in both homeland (Hong Kong) and heritage (Toronto) varieties of Cantonese.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.000 |
| Meta-epidemiology (narrow) | 0.003 | 0.004 |
| Meta-epidemiology (broad) | 0.005 | 0.001 |
| Bibliometrics | 0.005 | 0.005 |
| Science and technology studies | 0.001 | 0.002 |
| Scholarly communication | 0.001 | 0.016 |
| Open science | 0.005 | 0.003 |
| Research integrity | 0.003 | 0.007 |
| Insufficient payload (model declined to judge) | 0.008 | 0.002 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; both teacher heads agree on what is shown here.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".