Evaluating the accuracy of forced alignment across Mandarin varieties
Bibliographic record
Abstract
Forced alignment is widely used in phonetics to align transcripts with acoustic signals. These tools are trained on specific language varieties; it is unclear if they generalize to others. Previous research on English by MacKenzie and Turton (2020) finds good agreement between automated and human alignments for varieties that differ from the training variety. Such evaluation has only been carried out for English. We evaluate the level of human-aligner agreement on four Mandarin varieties (Canto, Shanghai, Beijing, and Tianjin). For each variety, two recordings from the HUB5 Corpus (LDC 1998) were aligned manually and by the Montreal Forced Aligner [McAuliffe et al. (2017)] using acoustic models trained on Beijing, Wuhan, and Hekou Mandarin [Schultz (2002)]. We find strong agreement between human and machine-aligned phone boundaries, with 17 ms as the median onset displacement. A mixed model identifies little variation across varieties or according to speech rate, but significant interindividual variation. Notably, despite the generally close agreement between the machine and human alignments, for two of the speakers, more than 10% of the alignments are displaced by over 100 ms. In sum, the Mandarin forced-aligner yields reliable alignments for out-of-training varieties, but manual checking of the results is still crucial.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.001 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".