Towards convergence of methods for speech and sign segmentation
Bibliographic record
Abstract
Signed languages, like spoken languages, combine sequences of arbitrary, finite, distinctive units in a continuous stream of movement (Stokoe 1960). Also like speech, determining where a particular sign begins and where it ends in a signing stream is not an easy task. Some researchers have established as the beginning of a sign the moment the hand is placed at the location where a certain sign is going to be initially or entirely produced and as its end the moment when the hand starts moving back to the rest position or to the location of the following sign (Crasborn & Zwitserlood 2008, Johnston 2009, Johnson & Liddell 2011). By doing so, these researchers leave transitional movements out of the limits of a sign. An alternative view claims that transitions should be partially or entirely regarded as part of a sign. Supporters base this view on the observations that (1) some articulatory features of a sign are visible even before or still after a sign is produced and (2) perceivers are able to guess signs solely drawing on information conveyed during transitions (Kita et al 2006, Bressem 2011, Jantunen 2010, 2013, 2015). The present study uses video data of Brazilian Sign Language (Libras; Xavier 2014) to critically evaluate the criteria traditionally used to delimit lexical items in the sign stream. Results indicate that methods used by speech researchers to delimit units in the speech stream are likely to be a good fit for delimiting units in the sign stream as well. Implications for speech and signed motor control will be discussed.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.000 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.000 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".