Using criterio voice familiarity to augment the accuracy of speaker identification in voice lineups
Bibliographic record
Abstract
Voice familiarity is a principal factor underlying the apparent superiority of human-based vs machine-based speaker identification. Our study evaluates the effects of voice familiarity on speaker identification in voice lineups by using a Familiarity Index that considers (1) recency (the time of last spoken contact), (2) duration of spoken contact, and (3) frequency of spoken contact. Three separate voice-lineups were designed each containing ten male voices with one target voice that was more or less familiar to individual listeners (13 per lineup, n = 39 listeners in all). The stimuli consisted in several verbal expressions varying in length, all of which reflected a similar dialect and the voices presented a similar speaking fundamental frequency to within one semitone. The main results showed high rates of correct target voice identification across lineups (>99%) when listeners were presented with voices that were highly familiar in terms of all three indices of recency, duration of contact, and frequency of contact. Secondary results showed that the length of the verbal stimuli had little impact on identification rates beyond a four-syllable string.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.002 | 0.001 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.001 |
| Science and technology studies | 0.000 | 0.000 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.002 | 0.000 |
| Research integrity | 0.000 | 0.000 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".