Representation matters: Developing a Canadian BIPOC composers dataset for music collection evaluation and development
Bibliographic record
Abstract
The music profession and industry, especially in traditions of western art music, is marked noticeably by a lack of compositions by Black, Indigenous and People of Colour (BIPOC). This lack of representation is just one of the many effects of generations of colonization, systematic exclusion, bias, and racism. There are numerous consequences to curating music library collections that continue to exclude BIPOC composers and artists, most notably giving the impression that such individuals do not exist or that their works are not worthy of inclusion. This potentially leads to a ripple effect whereby it becomes harder to program music by BIPOC composers, teach it, and write about it. This presentation describes the process and development of a dataset of BIPOC Composers with a connection to Canada, a project undertaken at the University of Saskatchewan (Treaty Six Territory and Homeland of the Metis, Saskatoon SK, Canada) through the work of the University of British Columbia School of Information Professional Experience Program. This project aimed to identify composers who identify as BIPOC and Canadian, or who identify as BIPOC and are based in what is now known as Canada. The project’s end goal was evaluating BIPOC representation in the University of Saskatchewan Libraries music collections, and ultimately filling collection gaps where needed. The dataset primarily serves as a tool for internal collection assessment but will be published and preserved in an open format for others who may be doing similar work. We will discuss the challenges associated with identifying BIPOC composers, especially in a Canadian context, and explore some of the ethical considerations when attempting to classify professionals using markers such as ethnicity or nationality.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.004 | 0.013 |
| Meta-epidemiology (narrow) | 0.003 | 0.001 |
| Meta-epidemiology (broad) | 0.001 | 0.002 |
| Bibliometrics | 0.010 | 0.008 |
| Science and technology studies | 0.005 | 0.001 |
| Scholarly communication | 0.004 | 0.002 |
| Open science | 0.005 | 0.004 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.026 | 0.029 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".