Methodological Trade-offs for Dual-purpose Phonetic Fieldwork
Bibliographic record
Abstract
Ultrasound overlay videos involve the superposition of ultrasound imaging of the tongue onto facial profile videos in order to serve as instructional materials (Abel et al., 2015). Bliss et al. (2016) used this technique to develop instructional and cultural materials for Indigenous communities by creating custom overlay videos of community members which highlight difficult sound contrasts in the languages for learners. Building on this work, this paper reports on the creation of this type of video for Hän, a Dene/Athabaskan language of Eagle, Alaska and Dawson City, Yukon with 6-7 native speakers remaining. In the paper, we explore the challenges behind a new possibility of collecting the ultrasound data recorded for these instructional videos to serve a dual-purpose: instructional/cultural and linguistic/scientific. Some acoustic work has been done on Hän (Manker 2012), but never articulatory. As Hän is known for its large phonemic inventory, being tied for first as the language with the most affricates and containing a 5-6 way contrast in the coronal region, articulatory work on the language is of interest for phonetic and phonological theory. However, ultrasound work has quite strict methodological standards, which can be impractical to include in many field situations, such as precise head and probe stabilization, and fully controlled phonological environments and speaker groups. These standards by necessity and design could not be fully adhered to for the Hän recordings. But, despite the methodological limitations involved in dual-purpose fieldwork, we argue that it is important to consider the possibility of drawing linguistic insights from data collected for instructional purposes. Otherwise, these insights simply wouldn’t exist as work with these communities is limited. References J. Abel, B. Allen, S. Burton, M. Kazama, M. Noguchi, A. Tsuda, N. Yamane, and B. Gick. Ultrasound-Enhanced Multimodal Approaches to Pronunciation Teaching and Learning. Canadian Acoustics, 43(3). 124-125, 2015. Bliss, H., Burton, S., and Gick, B. (2016). Ultrasound Overlay Videos and Their Application in Indigenous Langauge Learning and Revitalization. Canadian Acoustics, 44(3). Manker, Jonathan. (2012). An Acoustic Study of Stem Prominence in Hän Athabaskan. Master’s thesis, University of Alaska Fairbanks.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.147 | 0.253 |
| Meta-epidemiology (narrow) | 0.002 | 0.002 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.003 | 0.002 |
| Science and technology studies | 0.004 | 0.006 |
| Scholarly communication | 0.004 | 0.004 |
| Open science | 0.005 | 0.008 |
| Research integrity | 0.003 | 0.004 |
| Insufficient payload (model declined to judge) | 0.020 | 0.006 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".