Acquiring, Preserving, and Exhibiting a Comprehensive Collection of Vocal Music Recordings from Early- to Mid-Twentieth Century
Bibliographic record
Abstract
The Stratton-Clarke collection consists of approximately 200 linear feet of 78 and 33 1/3 rpm records, and thousands of digitized recordings that represents a comprehensive history of early twentieth century recorded Western sound, specifically opera -- its artists, roles, and early legacy from 78 rpm to early long play records. Along with someephemera and several pieces of historic playback equipment, a large financial gift will offset the costs of processing, preserving and providing access to the various formats represented in the collection. As the largest music research collection in Canada, the University of Toronto Music Library is fortunate to have the capacity to manage a donation of this magnitude. Each of our four authors has an important role to play to make the project a success. In this article we present a history and background of John Stratton, Stephen Clarke, and the collection itself, and document the many facets of a library taking on a donation of this size: donor relations and collaboration with the University’s advancement team and other stakeholders; the project management involved in making space and designing workflow for cataloguing, processing, and storage; archival description of the 78s and ephemera; preservation of the digital objects and digitization strategies for the analog recordings; the challenges and opportunities of working with large financial gifts; teamwork and managing students; and future plans for physical and online exhibitions of the collection.
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame machine prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. The Gemma side is a direct model label for every work in the frame, read from the title-only record. The Codex side is a classifier learned from the 10,348 direct Codex labels and calibrated to design-weighted sample rates; fields without enough sample support carry no Codex call. Candidate is the union of the two sides; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels.
Distilled classifier scores by category (both heads)
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.006 | 0.010 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.000 | 0.000 |
| Bibliometrics | 0.007 | 0.005 |
| Science and technology studies | 0.004 | 0.007 |
| Scholarly communication | 0.008 | 0.005 |
| Open science | 0.002 | 0.005 |
| Research integrity | 0.002 | 0.002 |
| Insufficient payload (model declined to judge) | 0.011 | 0.005 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one source (direct Gemma or distilled Codex), not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".