Canadian Dataverse Admin Survey - Individual (2023)
Bibliographic record
Abstract
The Canadian Dataverse Admin Survey - Individual (2023) was conducted to gain insight into the growing community of Canadian Dataverse Collection Administrators—librarians or other institutional staff members who manage an institutional Dataverse collection (or repository) and provide support to the members of their designated local research community. The Principal Investigators (PIs) of the study are Meghan Goodchild, PhD (Queen’s University / Scholars Portal) and Alisa Rod, PhD (McGill University), who lead the “Canadian Dataverse Administrators Survey Working Group,” a subgroup of the Dataverse North Expert Group of the Digital Research Alliance of Canada. Working Group members include: Shahira Khair (University of Victoria), Danica Evering (McMaster University), Alexander Jerabek (Université du Québec à Montréal), Tara Stieglitz (MacEwan University), Lacey Cain (Carleton University), and Lina Marie Harper (Digital Research Alliance of Canada), with support from John Huck (University of Alberta) and Amber Leahey (Scholars Portal). This dataset contains a README file (TXT), documentation (PDFs), the de-identified quantiative dataset (CSV), and R markdown notebook scripts to produce the figures (RMD). Le sondage sur les administrateur(trice)s du Dataverse canadien – Individus (2023) a été menée pour étudier sur la communauté grandissante des administrateur(trice)s du Dataverse canadien, soit les bibliothécaires et autres membres du personnel qui gèrent une collection Dataverse institutionnelle (ou un dépôt) et fournissent du soutien aux membres de leur communauté de recherche locale désignée. Les chercheuses principales du sondage sur les administrateur(trice)s du Dataverse canadien sont Meghan Goodchild, Ph. D. (Université Queen’s/Scholars Portal) et Alisa Rod, Ph. D. (Université McGill), qui dirigent le « groupe de travail sur le sondage sur les administrateur(trice)s du Dataverse canadien », un sous-groupe du groupe d’experts sur Dataverse Nord de l’Alliance de recherche numérique du Canada. Les membres du groupe de travail sont Lacey Cain (Université Carleton), Danica Evering (Université McMaster), Lina Marie Harper (Alliance de recherche numérique du Canada), Alexander Jerabek (Université du Québec à Montréal), Shahira Khair (Université de Victoria) et Tara Stieglitz (Université MacEwan), avec le soutien de John Huck (Université de l’Alberta) et d’Amber Leahey (Scholars Portal). Cet ensemble de données contient un fichier README (TXT), de la documentation (PDF), l'ensemble de données quantitatives anonymisées (CSV) et des scripts de carnets de démarquage R pour produire les chiffres (RMD).
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.008 | 0.028 |
| Meta-epidemiology (narrow) | 0.001 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.006 | 0.017 |
| Science and technology studies | 0.005 | 0.001 |
| Scholarly communication | 0.005 | 0.002 |
| Open science | 0.002 | 0.003 |
| Research integrity | 0.001 | 0.002 |
| Insufficient payload (model declined to judge) | 0.157 | 0.056 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".