Using criterio voice familiarity to augment the accuracy of speaker identification in voice lineups
Bibliographic record
Abstract
Voice familiarity is a principal factor underlying the apparent superiority of human-based vs machine-based speaker identification. Our study evaluates the effects of voice familiarity on speaker identification in voice lineups by using a Familiarity Index that considers (1) recency (the time of last spoken contact), (2) duration of spoken contact, and (3) frequency of spoken contact. Three separate voice-lineups were designed each containing ten male voices with one target voice that was more or less familiar to individual listeners (13 per lineup, n = 39 listeners in all). The stimuli consisted in several verbal expressions varying in length, all of which reflected a similar dialect and the voices presented a similar speaking fundamental frequency to within one semitone. The main results showed high rates of correct target voice identification across lineups (>99%) when listeners were presented with voices that were highly familiar in terms of all three indices of recency, duration of contact, and frequency of contact. Secondary results showed that the length of the verbal stimuli had little impact on identification rates beyond a four-syllable string.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.003 | 0.023 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.000 | 0.001 |
| Scholarly communication | 0.001 | 0.001 |
| Open science | 0.000 | 0.001 |
| Research integrity | 0.001 | 0.000 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".