Tonal alternations in attributive constructions in Mwaghavul
Bibliographic record
Abstract
Mwaghavul is an underdocumented Chadic language spoken in Plateau State, Nigeria, by approximately 150,000 people (Blench 2011). Mwaghavul has tonal lowering in associative constructions, where the first nominal in the construction surfaces with low tone, regardless of its tone in isolation (Arokoyo & Fwangwar 2019). However, tonal lowering is not fully predictable, as some high tone nominals surface as mid tone in associative constructions, instead of low. Number of syllables, vowel length and quality are not consistent predictors, as there are minimal pairs for high tone alternations. We investigate the phonetics of these high tones, to determine whether two phonetically distinct high tones have been incorrectly documented as one, or whether one phonetic high tone has two phonological behaviours. The f0 of 77 tokens in isolation and 561 tokens in associative constructions was extracted at 10 points using Prosody Pro (Xu 2013). In isolation, high tones that become mid in associative are visually distinct from those that become L, with approximately 15-20Hz difference throughout the tone duration. Linear mixed effects models confirm this difference is statistically significant. The presence of separate high and superhigh tones in Mwaghavul indicates that the phonetic implementation of the floating low tone is realized differently depending on the pitch of the original tone. This suggests that the original tone is not deleted, but rather dissociated and present, affecting the realization of the tonomorpheme in an unusual pattern that is not commonly attested.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.003 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".