MétaCan
Menu
Back to cohort
Record W4393662717 · doi:10.5281/zenodo.8318811

Interruption Audio & Transcript: Derived from Group Affect and Performance Dataset

2023· dataset· en· W4393662717 on OpenAlexaboutno aff
D. John Doyle, Ovidiu Şerban

Bibliographic record

VenueZenodo (CERN European Organization for Nuclear Research) · 2023
Typedataset
Languageen
FieldPsychology
TopicTeam Dynamics and Performance
Canadian institutionsnot available
Fundersnot available
KeywordsAffect (linguistics)Group (periodic table)CommunicationSpeech recognitionComputer sciencePsychologyChemistry

Abstract

fetched live from OpenAlex

Licensing This dataset is adapted from the Group Affect and Performance dataset which is released under the Creative Commons Attribution-NonCommercial 4.0 International (CC BY-NC 4.0) license. https://creativecommons.org/licenses/by-nc/4.0/ Description This dataset contains the audio files containing manually annotated cases of overlapped utterances, classified into True Interruptions and False Interruptions. It is derived from the Group Affect and Performance dataset created by the University of the Fraser Valley, Canada. Original conversation transcripts and audio files have been supplied for context. The Group Affect and Performance dataset provides a rich source of interruptions and overlapped utterances in general, yielding 200 True Interruptions from 355 instances of overlapped utterances in the 14 Group meetings which were annotated. Structure This dataset is structured into three parts: 1. data.json contains a list of all instances of overlapped utterances, classified into ‘interruption’ and ‘non-interruption’ corresponding to True and False Interruptions respectively. Each instance is uniquely identified by the Group in which it occurred, the speaker and the starting time of the utterance. 2. The 'audio' directory contains the audio of each instance of overlapped utterances corresponding to those found in data.json. The naming convention of the files is as such: ‘Group [group number]: [utterance start time] - [utterance end time].wav’. 3. Also included is a copy of the original dataset which includes the full audio and transcript. This allows the full meeting to be heard and any context for interruptions to be evaluated. Note that directories 2. and 3. can be accessed by unzipping audio-and-transcripts.zip. Data Collection Protocol Of paramount importance to our process are the definitions of an overlapped utterance and a True Interruption. A False Interruption is simply an overlapped utterance which is not a True Interruption. These definitions directly impact the dataset; for overlapped utterance it informs which data points are included in our dataset and for True Interruption it informs the classes assigned to each sample. In defining an overlapped utterance, our primary aim is to create an overarching class encompassing interruptions and all instances that could be deemed a True Interruption. For this reason, we omit cases where the timing misplaced speech and early-onset responses. An overlapped utterance is defined as an instance where one interlocutor provides speech or noise during another interlocutor’s speech, creating an overlap that may be deemed a possible interruption when considering its timing alone. For this reason we omit cases of where the timing indicates misplaced speech or early-onset responses. Our definition of True Interruption is an instance where an interrupting party intentionally attempts to take over a turn of the conversation from an interruptee and, in doing so, creates an overlap in speech. As previously mentioned, due to the ‘intent’ part of this definition, we avoid cases of misplaced speech and early-onset responses. The former is enforced by not considering cases of overlapped speech which begin within 300ms of each other since this is an estimate for the average human reaction time of articulating a vowel in response to a speech stimuli. The latter is enforced by not considering speech starting within the last 10% of first utterance in the overlapped speech. Note that this approach fails to filter out all cases of misplaced speech, so we manually remove the remaining instances. Methodology Three main steps were taken to produce this dataset: 1. Parsing the transcripts for cases of overlapping speech 2. Manually annotating these cases per our protocol and adding them to data.json 3. Extracting audio samples from data.json and adding them to the audio folder If you use this dataset, please cite the following paper: Doyle, D.; Şerban, O. Interruption Audio & Transcript: Derived from Group Affect and Performance Dataset. Data 2024, 9, 104. https://doi.org/10.3390/data9090104 @article{data9090104, AUTHOR = {Doyle, Daniel and Şerban, Ovidiu}, TITLE = {Interruption Audio & Transcript: Derived from Group Affect and Performance Dataset}, JOURNAL = {Data}, VOLUME = {9}, YEAR = {2024}, NUMBER = {9}, ARTICLE-NUMBER = {104}, URL = {https://www.mdpi.com/2306-5729/9/9/104}, ISSN = {2306-5729}, DOI = {10.3390/data9090104} }

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame machine prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.

metaresearch head score (Codex)0.001
metaresearch head score (Gemma)0.005
Version: metacan-v3-hybrid-931329e0061cValidation status: machine_predicted_unvalidated
Candidate categoriesnone
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Dataset · Consensus signal: Dataset
Teacher disagreement score0.030
Threshold uncertainty score0.099

Distilled classifier scores by category (both heads)

CategoryCodexGemma
Metaresearch0.0010.005
Meta-epidemiology (narrow)0.0020.000
Meta-epidemiology (broad)0.0010.001
Bibliometrics0.0030.002
Science and technology studies0.0010.000
Scholarly communication0.0020.001
Open science0.0020.003
Research integrity0.0020.002
Insufficient payload (model declined to judge)0.0300.048

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.061
GPT teacher head0.304
Teacher spread0.244 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.

The models applied no category: nothing in the taxonomy fit this work.
Study designNot applicable
Domainnot available
GenreDataset

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2023
Admission routes1
Has abstractyes

Explore more

Same venueZenodo (CERN European Organization for Nuclear Research)Same topicTeam Dynamics and PerformanceFrench-language works237,207