Phonetic correlates of phonological quantity of Yakut
Bibliographic record
Abstract
We investigated vowel quantity in Yakut (Sakha), a Turkic language spoken in Siberia by over 400,000 speakers in the Republic of Sakha (Yakutia) in the Russian Federation. Yakut is a quantity language; all vowel and consonant phonemes have short and long contrastive counterparts. The study aims at revealing acoustic characteristics of the binary quantity distinction in vowels. We used two sets of data: (1) A female native Yakut speaker read a 200-word list containing disyllabic nouns and verbs with four different combinations of vowel length in the two syllables (short–short, short–long, long–short, and long–long) and a list of 50 minimal pairs differing only in vowel length; (2) Spontaneous speech data from 9 female native Yakut speakers (aged 19–77), 200 words with short vowels and 200 words with long vowels, were extracted for analysis. Acoustic measurements of the short and long vowels’ f0-values, duration and intensity were done. Mixed-effects models showed a significant durational difference between long and short vowels for both data sets. However, the preliminary results indicated that, unlike in quantity languages like Finnish and Estonian, there was no consistent effect of f0 as the phonetic correlate in Yakut vowel quantity distinction.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.002 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.001 | 0.001 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.001 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".