MétaCan
Menu
Back to cohort
Record W4393728959 · doi:10.5281/zenodo.4785016

CSD: Children's Song Dataset for Singing Voice Research

2021· dataset· en· W4393728959 on OpenAlexaff
Soonbeom Choi, Won‐Il Kim, Saebyul Park, Sangeon Yong, Juhan Nam

Bibliographic record

VenueZenodo (CERN European Organization for Nuclear Research) · 2021
Typedataset
Languageen
FieldComputer Science
TopicMusic and Audio Processing
Canadian institutionsKootenay Association for Science & Technology
Fundersnot available
KeywordsSingingSpeech recognitionAudiologyCommunicationPsychologyComputer scienceAcousticsMedicinePhysics

Abstract

fetched live from OpenAlex

Children's Song Dataset is open source dataset for singing voice research. This dataset contains 50 Korean and 50 English songs sung by one Korean female professional pop singer. Each song is recorded in two separate keys resulting in a total of 200 audio recordings. Each audio recording is paired with a MIDI transcription and lyrics annotations in both grapheme-level and phoneme-level. <strong>Dataset Structure</strong> The entire data splits into Korean and English and each language splits into 'wav', 'mid', 'lyric', 'txt' and 'csv' folders. Each song has the identical file name for each format. Each format represents following information. Additional information like original song name, tempo and time signature for each song can be found in 'metadata.json'. 'wav': Vocal recordings in 44.1kHz 16bit wav format 'mid': Score information in MIDI format 'lyric': Lyric information in grapheme-level 'txt': Lyric information in syllable and phoneme-level 'csv': Note onsets and offsets and syllable timings in comma-separated value (CSV) format <strong>Vocal Recording</strong> While recording vocals, the singer sang along with the background music tracks. She deliberately rendered the singing in a “plain” style refraining from expressive singing skills. The recording took place in a dedicated soundproof room. Singer recorded three to four takes for each song and the best parts are combined into a single audio track. Two identical songs with different keys are discriminated by character 'a' and 'b' at the end of a filename. <strong>MIDI Transcription</strong> The MIDI data consists of monophonic notes. Each note contains onset and offset times which were manually fine-tuned along with the corresponding syllable. MIDI notes do not include any expression data or control change messages because those parameters can be ambiguous to define for singing voice. Singing voice is an highly expressive sound and it is hard to define precise onset timings and pitches. We assumed one syllable matches with one MIDI note and made the following criteria to represent various expressions in singing voice. A piano sound is used as a reference tone for the annotated MIDI to ensure the alignment with vocal. The rising pitch at the beginning of a note is included within a single note. The end of syllable is treated as the offset of a note. The breathing sound during short pauses is not treated as note onset or offset. Vibrations are treated as a single sustaining note. If a syllable is rendered with several different pitches, we annotated them as separate notes. <strong>Lyric Annotation</strong> Text files in the 'lyric' folder contains raw text for corresponding audio and the 'txt' folder contains phoneme-level lyric representation. The phoneme-level lyric representation is annotated in a special text format. Phonemes in a syllable are tied with underbar('_') and syllables are separated with space(' '). Each phonemes are annotated based on the international phonetic alphabet (IPA) and romanized symbols are used to annotate IPA symbols. You can find romanized IPA symbols and more detailed information in this repository. <strong>License</strong> This dataset was created by the KAIST Music and Audio Computing Lab under Industrial Technology Innovation Program (No. 10080667, Development of conversational speech synthesis technology to express emotion and personality of robots through sound source diversification) supported support by the Ministry of Trade, Industry &amp; Energy (MOTIE, Korea). CSD is released under Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International (CC BY-NC-SA 4.0).<strong> It is provided primarily for research purposes and it is prohibited to be used for commercial purposes. When sharing your result based on CSD, any act that defames the original singer is strictly prohibited.</strong> For more details, we refer to the following publication. We would highly appreciate if publications partly based on CSD quote the following publication: Choi, S., Kim, W., Park, S., Yong, S., &amp; Nam, J. (2020). Children’s Song Dataset for Singing Voice Research. 21th International Society for Music Information Retrieval Conference (ISMIR). We are interested in knowing if you find CSD useful. If you use CSD please email us at kaist.mac@gmail.com and tell us about your research.

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.003
metaresearch head score (Gemma)0.002
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesMeta-epidemiology (narrow), Science and technology studies, Scholarly communication, Open science, Insufficient payload (model declined to judge)
Consensus categoriesInsufficient payload (model declined to judge)
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Dataset · Consensus signal: Dataset
Teacher disagreement score0.048
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0030.002
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.000
Bibliometrics0.0010.002
Science and technology studies0.0060.000
Scholarly communication0.0060.001
Open science0.0060.007
Research integrity0.0000.001
Insufficient payload (model declined to judge)0.0030.005

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.059
GPT teacher head0.300
Teacher spread0.241 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; both teacher heads agree on what is shown here.

Study designNot applicable
Domainnot available
GenreDataset

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2021
Admission routes1
Has abstractyes

Explore more

Same venueZenodo (CERN European Organization for Nuclear Research)Same topicMusic and Audio ProcessingFrench-language works237,207