Groupsourcing Folklore Sound Files: Involving the Community in Research
Bibliographic record
Abstract
AbstractAbstractDigital technologies make possible new ways of managing folklore field recordings. Two programmers, a graduate student, and the author developed a database that allows the user to go directly to the point in a sound file where a particular topic is discussed. This is a research tool and the task here was to create a modified site for the general public. The technique used was crowdsourcing, asking the public to transcribe and translate songs, stories, and accounts of belief. The project revealed how heritage issues affect public participation. People who expressed initial enthusiasm were reluctant to participate because they were timid about their language knowledge. Paradoxically, formal instruction leads to the timidity we observed. People who did contribute to our project transcribed and translated songs only. Language is retained in song even as it is lost elsewhere. Songs are also familiar material, associated with the past. Contributions driven by interest in new material directly from Ukraine did not materialize. A romanticized image of the past and suspicions that Ukraine has been Sovietized and Russified encourage preservation of the old and work against interest in the new.RésuméLes technologies numériques rendent possibles des méthodes de gestion nouvelles des enregistrements de folklore, faits dans le cadre d'un travail en terrain. Deux programmeurs, une étudiante graduée, ainsi que l'auteur elle-même, ont développé une base de données permettant à l'utilisateur d'accéder directement à un sujet précis dans un fichier sonore. Le but de ce moteur de recherche était la création d'un site web, modifié expressément pour les besoins du grand public. Nous avons utilisé la méthode de collaborât externe (crowdsourcing) en demandant au public de transcrire et traduire des chansons, des histoires et des dires relatant certaines croyances populaires. Le projet a révélé la façon dont les questions du patrimoine culturel affectent la participation du grand public. Les personnes ayant exprimé un intérêt initial pour le projet furent réticentes à y participer à cause de leur connaissance de la langue. Paradoxalement, leur instruction formelle a mené à cette timidité. Les personnes ayant participé au projet ont transcrit et traduit uniquement des chansons. La langue est maintenue dans les chansons même si elle est perdue ailleurs. Les chansons représentent aussi du matériel déjà connu et associé avec le passé. Des participants intéressés par de nouvelles données provenant directement d'Ukraine ne se sont pas encore matérialisés. Une image romanesque du passé et la méfiance quant à une Ukraine soviétisée et russifiée, semblent avoir encouragé une préservation de l'ancien et aller à l'encontre d'un intérêt pour le nouveau.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.000 | 0.000 |
| Science and technology studies | 0.003 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.001 | 0.000 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.002 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".