MétaCan
Menu
Back to cohort
Record W2325003127 · doi:10.1525/jams.2016.69.1.255

Review: Sheet Music Consortium

2016· article· en· W2325003127 on OpenAlexaboutno aff
Judy Tsou

Bibliographic record

VenueJournal of the American Musicological Society · 2016
Typearticle
Languageen
FieldArts and Humanities
TopicDigital and Traditional Archives Management
Canadian institutionsnot available
Fundersnot available
KeywordsIconCitationScholarshipHistoryLibrary scienceArt historyArtComputer sciencePolitical scienceLaw

Abstract

fetched live from OpenAlex

Historically, popular sheet music has been viewed as the poor cousin of regular music scores because it was not considered relevant to serious music scholarship. But there is currently serious work under way in this field, raising many interesting questions in relation to the publishing history, reception, collecting, even artwork and other decorative imagery of the sheet music that once lurked in piano benches, then slipped into attics, and now sits quietly on library shelves. Projects such as the Sheet Music Consortium (SMC) can advance this work considerably, permitting scholars to search in several collections at once with new precision. The project is a collaboration among libraries to build “an open collection of digitized sheet music using the Open Archives Initiative Protocol for Metadata Harvesting,” or OAI-PMH.1 Metadata, technically, is data about data. Practically speaking, the metadata, which amounts to the indexing of names, titles, publishers, and so on, allows easy retrieval of the sheet music through these indices. As we will see, this metadata can relate to almost any aspect of the score as an object, the music and lyrics it contains, even the artwork on the cover. It is therefore essential to have good metadata in any digital project. The parameter for “sheet music,” as used here and at large, is solely based on the physical format of the score: generally, the music is in loose sheets or folios, each piece being from one to ten pages long. Most sheet music consists of popular songs with piano accompaniment, even though some classical music also appears in this format. In this project, both genres are present, popular music forming the bulk of the collection.The original portal for the SMC was launched in 2003, after two years of planning, by the four founding institutions: the University of California, Los Angeles (UCLA), Indiana University, Johns Hopkins University, and the Library of Congress. This phase of the project was supported by a grant from the Mellon Foundation and hosted by the UCLA Digital Library Program. Each institution brought slightly different expertise and resources to the project, and it was important to have the institutions’ support to ensure long-term commitment. Between 2007 and 2008 a federal planning grant from the Institute for Museum and Library Services (IMLS) supported the consortium's work in publicizing the project and planning a second generation of the website, including the use of the Semantic Web and Linked Open Data, which I will discuss later. The publicity during the second grant resulted in the participation of many more institutions, expanding from 78,708 items in the original four collections in 2002 to 226,914 items in twenty-two collections in February 2013. At the time of writing, the “Featured Repositories” page shows thirty-nine collections with 358,922 items,2 and includes collections from Maine to Texas and from Washington, D.C., to Canberra, Australia.The topics of the songs in the combined collection vary, from love, to ethnic “comic” songs, to specific events such as World War I or the World's Fair. Many of the individual collections also include songs with regional themes. The Auburn University and Duke University collections, for instance, contain Southern-themed music that offers antebellum music and Confederate imprints. The collections at the Australian National Library and York University (Canada) add an international flavor to the geographic mix. Collectively, the collections offer music from the 1700s to the present, most of it dating from between 1850 and 1950, the heyday of sheet music as in-home entertainment. In addition, in the post–Civil War period a new method of publishing using the stereotype process allowed the mass production of sheet music. This is reflected in the number of items listed for each decade on the “Date Published” page (under “Browse”). Because the SMC includes many large collections of sheet music, it probably offers a good representative sample of the trend in sheet music publication during those decades.Another way to preview the scope of the collection is by perusing the various tabs on the “Browse” page: “Title,” “Subject,” and “Name,” all with convenient alphabetical jump lists for direct access to a specific letter. Examining the items under “A” for the “Subject” list, we see names as subjects and portraits of some of these named subjects, Aboriginal Australian songs, sixty-six songs on Abraham Lincoln alone, and songs on abused children and abused women. This index also readily shows the inconsistent use of standardized terms in subject entries among catalogers. For example, there are two forms of the word “absence”: “Absence” and “Absences.” The inconsistency shows that not every cataloger was following the recommended Library of Congress authority files and thesaurus,3 as advised in the manual Cataloging Sheet Music.4 Thus, an understanding of the underlying controlled vocabulary and its limitations will aid the use of the site.Even though the stated goal of the project is to build a collective digitized sheet music collection, it is important to note that the SMC harvests only the metadata.5 Any existing digital images remain at the contributing institution with links from the website, and not all institutions have digitized the materials whose metadata they have made available. The lack of some digital scores is in large part due to the copyright law restricting public domain scores to pre-1923 publications;6 many individual collection websites state this restriction clearly. Another possible reason for the paucity of digital images is that most sheet music collections tend to be large, making digitization of every item unscalable. A case in point is the Duke University collection, where only about 3,000 of the 20,000 items in the collection are digitized. Duke chose to digitize only music published in America between 1850 and 1920 in order to represent “a wide variety of music types including bel canto, minstrel songs, protest songs, sentimental songs, patriotic and political songs, plantation songs, Civil War songs, spirituals, dance music, songs from vaudeville and musicals, ‘Tin Pan Alley’ songs, and songs from World War I.”7 The goal is to highlight some of the unique items, a strategy that is common in large collections.Like digitization, the cataloging of each title is a daunting task on account of the large numbers involved. Many libraries either catalog them at the collection level (one bibliographic record for the entire collection) or leave them uncataloged. Collection-level cataloging does not yield much useful data for the scholar or general user, as not even individual song titles can be retrieved. It may be useful if the collection has a focus, allowing the cataloger to name it, for example, “Jane Doe's Civil War Music.” Otherwise collection-level cataloging is merely to let the user know of the collection's presence in a particular library.Fortunately for users of the SMC, each of the collections represented is richly cataloged at the item level. Such work is expensive and time-consuming, but of great benefit to scholarly work. Special guidelines for cataloging sheet music were first developed by Brown University in the early 1980s.8 These guidelines are especially useful in that sheet music access is often different from that of a regular score. For example, many scholars study the images on the covers of sheet music and so keyword descriptions of these images, such as “gate,” “moon,” and “girl,” are important; these are retrievable in a subject search. The SMC cataloging depends on a series of evolving standards, from the old MARC cataloging fields developed in the 1960s,9 to Brown's sheet music guidelines, to the recently updated Dublin Core Metadata, a set of standardized vocabulary terms for cataloging physical as well as digital objects.10 Only one field, that of title, is required in the SMC metadata, but fortunately institutions generally supply all the other basic fields—creators, publishers, and dates—while many collections provide much more detail that is of great interest to scholars.The opening screen provides a single search box by which to search all thirty-nine collections at the same time. In addition, there are options to search in the title, name, subject, place, or publisher indices. Users can also limit the search to digitized music only, a useful feature, because, as mentioned above, only a portion of the holdings is available in digital facsimiles. But the metadata provides powerful points of access and discovery for individual items and large-scale trends that might be impossible to measure working with individual collections. Users can specify as many of the indices as they wish, thanks to the “Advanced Search” menus, and add multiple fields for any of the larger categories of metadata (including places, publishers, authors, and so on). Just as it is important to understand the limitations of the underlying metadata that stands behind these sets, it is also important to understand the principles of Boolean searching when using the various layers of fields. Even though it is not specified, for instance, the basic search is not a limiting function;11 rather, it functions like a Boolean “or” rather than “and” operation. Thus, the more indices chosen, the greater number of search results. For example, if you put “New York” in the search box and click the “Titles” box you get 3,555 results; if you perform the same search on “Publisher” you get 52,063; if you search both “Titles” and “Publisher” you get 55,185, combining the two pools of results. (The discrepancy in the numbers is due to the fact that “New York” occurs in both title and publisher in some items.)In the “Advanced Search” option, however, a search on multiple fields is equivalent to a Boolean “and” search, generating fewer results as you add more indices through the “Add Row” button. For example, if you are looking for songs by Rodgers and Hammerstein you can search under “names” for “Richard Rodgers” and “Oscar Hammerstein” in two rows, and the results will all include both names. The fields that may be added are “keyword,” “title,” “names,” “subjects,” “place,” and “publisher” (see Figure 1).This is all quite promising, but the inability to customize the display makes it difficult to find specific kinds of information in the search results. In many other search interfaces, such as most library catalogs, results can be listed alphabetically by title or name, or they can be displayed in chronological or reverse chronological order, which allows the user to quickly find specific information or to survey trends. The current display of the search results shows no detectable logic in its arrangement. The results can nevertheless be filtered to a specific collection; local Durham residents, for example, may want to limit the search results to Duke University holdings only. Similarly, you can limit the publications to a date range. Like the “Basic Search” screen, the “Advanced Search” screen offers a filter for digitized scores. And you can combine any or all of the filters for the search results. For example, you can search for “love” in the keyword index and request digitized music from Duke University holdings published between 1910 and 1920.The search and display functions are generally adequate, though adding modern features such as faceted navigation and the ability to export search results would improve the site's usefulness. Faceted navigation groups search results into predetermined categories, such as resource types, publication dates, author, language, and so on, and simultaneously shows the number of results in each of these categories. In the SMC catalog it would be useful to have facets for creators, titles, subjects, publishers, publication dates, and collections. For example, if you were to perform a keyword search on “Chicago” you could easily see the number of items with “Chicago” as subject as opposed to place of publication. You could select the results for “Chicago” as a subject without having to perform a “limit” search, or manipulate the search results in a more focused way through these facets. Or if you were searching for the World's Columbian Exposition the search results would readily show which institutions have the most items, without the need to go to each institution's site. Such nimble searching is difficult with the current limited range of “limits.”Another feature that would be useful is the ability to directly export search results for later analysis and manipulation. Although this is not currently possible, the SMC team imagined a series of what they called “virtual collections” that might be used privately through individual accounts (not functional at present), shared with specific individuals, or publicly accessible. These functions are of merit as ways to promote individual scholarship, public discussion, and even classroom exploration of these materials. They do not, however, replace direct exporting of results.A private virtual collection is password protected; not even the title is visible to anyone other than the creator. The titles of shared collections, on the other hand, are visible to anyone, although the contents are password protected. The contents of the third type of virtual collection, the public collection, are viewable by anyone, without a password (see Figure 2). The creation of such a collection is through a “drag and drop” process, but the function is currently broken, making it impossible to test this out. The function apparently worked at one time, given that there are visible virtual collections on the site. The requester of a public collection is sent an e-mail notification with a link, but since this e-mail consists only of the collection's URL that links back to the web page, it is impossible to edit or manipulate the list. Traditional catalogs with export features send the entire list (either as an attachment or in the body of the e-mail) to the user, often allowing the user to apply standard styles in the process—styles such as those set by the Modern Languages Association. The user can then save the list for emendation and use. Despite the lack of flexibility, this virtual collection feature could still be very useful for researchers and teachers, when functional.The unclear status of the virtual collections notwithstanding, the SMC team, with its second grant, has developed a next-generation version of the resource that will ultimately take advantage of a concept called the Semantic Web, a common framework that allows data to be shared and reused across applications and community boundaries. In other words, the Semantic Web, with its markup, allows software to find things, understanding that, for example, “hamlet” might be a title, or a character, or a place. The new website, in addition to the new look, is more intuitive, and this may be the rationale for not having a help screen. It also provides the searching and browsing features mentioned above. Another improvement was the Metadata Mapping Tool,12 developed by Indiana University. The tool makes it “possible to upload, crosswalk and validate metadata from a variety of sources, including plain text files, Microsoft Excel spreadsheets (.xls or .xlsx), or dBase files (.dbf).”13Also in line with the Semantic Web, the group initiated the Linked Open Data (LOD) project to allow for deeper exploration of the metadata. The SMC started the pilot in June 2013, focusing on publishers. The rationale behind exposing the publisher data is to provide information on American music publishing history “in a [new] dimension.” The website continues,In order to expose the data, publisher names must be normalized across all collections within the SMC. In addition, each publisher will receive a permanent URI (Uniform Resource Identifier) with unique identifiers, and each of these IDs will have an associated permanent URL and a link to an RDF (Resource Description Framework) file, which facilitates data exchange and mashups. The website lists fourteen publisher records that are in the process of being enhanced. Only one publisher, however, Kalmar, Puck & Abrahams, is completed with the requisite RDF and XML.15 In these RDFs the publisher's address has a URL linked to the GeoNames database, a dataset independent of the SMC project that provides standardized geographic information to which other resources can point. It includes semantic chains that allow one to map it to other datasets. I mapped this single publisher onto the GeoNames dataset and located Kalmar et al. in the city of New York on the map.16 Not much can be deduced from a single datum, but once all the publishers have been mapped it will be possible to use the data in combination with the GeoNames database to discover interesting geographic trends in American sheet music publications. Music scholars are just into the digital and all of could be linked to interesting to questions on a large with and other that the from a even as individual publishers, or are many to collective like where metadata is of images being in one the like the SMC does not need the large of that is required to And the new to be a that may be in the only the metadata is at the SMC is much more And with the data the the other hand, there are some to such a project. Because metadata is rather than its In addition, the sheet music collections may have been cataloged the SMC project. Thus, the data may have to be if it is to with the Dublin Core Because of these the of search results is For example, University to include a of detail in its with descriptions of the and including the lyrics of all the and the of a Most other institutions, do not include such the SMC be as a for lyrics or other other than the basic is the multiple that a user may have to during a single in order to digital scores at different institutions, as from the SMC portal to the individual in order to or digital Many of the for these individual are and easy to but some are The navigation for the digital images on the Johns Hopkins for example, is on an you have to go back to the original page to a second page of the same piece of there is no from one page to the as is of modern Another is the University of where it is not easy to the digital on the for the song to you are the of The of which is not the you might a box that allows you to find information within the Another is to click on a of to find the page number of the to get to the song you must click on the linked page This is a and process of exploration to have to go through merely to display a single but it is the one if one to use a website such as the the behind the Sheet Music Consortium are But in a there are to what can be in a series of UCLA and Indiana University are to the support of the metadata and for relevant But there is not at by with based at institutions such as the Library of where an such as the Music Consortium has long-term to the SMC website is still the most important resource for anyone sheet music across collections. The though in many point the in the and to a large resource of digitized scores. The of being to do most of the from is to having to to different to find the relevant music It would be if UCLA and Indiana University with the support of the Music Library could the good work of the SMC in the years to

Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.

How this classification was reachedexpand

Full frame distilled prediction

Teacher imitation

Not calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.

metaresearch head score (Codex)0.000
metaresearch head score (Gemma)0.000
Version: codex-gemma-dda1882f352aValidation status: machine_predicted_unvalidated
Candidate categoriesInsufficient payload (model declined to judge)
Consensus categoriesnone
DomainCandidate signal: none · Consensus signal: none
Study designCandidate signal: Not applicable · Consensus signal: Not applicable
GenreCandidate signal: Empirical · Consensus signal: Empirical
Teacher disagreement score0.313
Threshold uncertainty score1.000

Codex and Gemma teacher scores by category

CategoryCodexGemma
Metaresearch0.0000.000
Meta-epidemiology (narrow)0.0000.000
Meta-epidemiology (broad)0.0000.001
Bibliometrics0.0000.000
Science and technology studies0.0000.001
Scholarly communication0.0000.000
Open science0.0000.000
Research integrity0.0000.000
Insufficient payload (model declined to judge)0.0010.000

Machine scores (provisional)

The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.

Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.

Opus teacher head0.046
GPT teacher head0.233
Teacher spread0.187 · how far apart the two teachers sit on this one work
Validation statusscore_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from it

Classification

machine, unvalidated

Machine predicted; a candidate call from one teacher head, not a consensus.

Study designNot applicable
Domainnot available
GenreEmpirical

How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".

Quick stats

Citations0
Published2016
Admission routes1
Has abstractyes

Explore more

Same venueJournal of the American Musicological SocietySame topicDigital and Traditional Archives ManagementFrench-language works237,207