Music’s AI Problem, AI’s Music Problem
Bibliographic record
Abstract
For David Huron“Today’s AI is the worst the technology will ever be,” or so the popular quip goes.The claim is obviously true. The breakneck pace of technological developments recycles today’s state-of-the-art futurism into tomorrow’s outdated fossils with a shocking speed. Contemporary computational hardware and software are impossibly more advanced than they were even five years ago, let alone compared with the punch-card programs for which the term Artificial Intelligence (AI) was coined over seventy years ago. As the first quarter of the twenty-first century transitions into the second, we’re witnessing a volcano of innovations across the computational gamut, including programming, processing, and data collection, all accessible with a few clicks on a browser or swipes on a smartphone.But, from a larger social perspective, the claim’s truth value is a bit more ambiguous. On the one hand, advanced AI might be an agent of change that shuttles us into a bright future of AI assistants, tech-supported efficiencies, and new modes of creativity and learning. On the other, it might erode fundamental aspects of our culture, workforce, education, and minds. We seem to be standing on the brink of larger changes, and it’s hard to tell whether the path forward slopes upward or downward.In this essay, I’m going to outline some of the broader topics surrounding musical AI, and a few issues I see orbiting the technology. I’m going to focus particularly on generative musical AI, Artificial Intelligence that creates music, but I’ll certainly touch on many other related AI applications. Throughout, I’ll connect the specifics of the technology to larger issues in academic music studies, and, in so doing, I’ll point out a few landmarks emerging in the hazy road ahead.The essay divides into seven sections. My first two sections provide historical and technical backgrounds of musical AI, with a particular emphasis on neural networks and how computers can be trained to produce music. Then, I’ll outline a few ways that AI is used in contemporary musical studies and composition. My fourth section deals with difficulties that music provides to AI engineers, and my fifth section shows how these difficulties manifest in actual generated musical content. The sixth section considers some ways that music and musicians both are and are not vulnerable to AI replacement, and, finally, I end by discussing what aspects of contemporary AI I’m worried about, what I’m not worried about, and what I’m excited about.The term Artificial Intelligence was coined in 1955 by a group of researchers spending their summer at the Dartmouth Summer Research Project on Artificial Intelligence, where they studied how computers can simulate human reasoning, problem solving, abstraction, and learning, with learning being perhaps the most important fulcrum of this new technology.Humans learn how to do tasks through a combination of implicit and explicit instruction. We learn how to speak our native language both by being exposed to other people speaking (implicit learning) and by teachers instructing us (explicit learning). Of the two, implicit learning has always been the greater prize for computational engineers, and it’s not hard to see why. In implicit machine learning—or, just machine learning—a program learns how to do some task simply by observing data fed into its programming. Cognitive scientists can use machine learning to study the human mind by computationally modeling some learning process. Data scientists can use it to mine trends and predictions from huge piles of statistics. And, most relevant to the current topic, engineers can design programs that teach themselves to do any task just by observing and learning from data, thereby obviating any need for meticulously and painstakingly programming these tasks into the machine.Plenty of music researchers and composers would experiment with machine learning in their work, with Knud Jeppeson’s 1923 investigation of Palestrina’s counterpoint serving as an analog prolog.1 Interested in reconstructing Palestrina-style counterpoint with the highest fidelity to the original source, Jeppeson meticulously tallied how often various contrapuntal behaviors appear in Palestrina’s works, and used those statistics to fashion textbook rules for writing counterpoint in this style. The result was a proto-machine-learned model, with trends drawn from a dataset with pen and paper, and calculated using long division. After academic institutions began investing in computers in the 1950s, it didn’t take long for musicologists to explore the technology’s research potential. In 1957, Leonard Meyer theorized that computers could process a dataset of a composer’s music to statistically capture the norms and tendencies of their musical style, essentially moving Jeppeson’s pen-and-paper research methods to the digital language of a computer. In the coming years, the likes of Joseph E. Youngblood, Edgar Coons, David Kraehenbuehl, Gift Siromoney, and K. R. Rajagopalan would implement Meyer’s theories with basic machine-learning models.2 In an early attempt to generate music using these norms and tendencies, the composer and engineer Raymond Scott programmed a vast analog machine with popular-music rhythms and melodic licks in the 1960s, and sold his machine to Motown records to provide their songwriters with musical ideas and inspiration.3 In 1981, the composer David Cope would create a machine-learning system to grapple with his own writer’s block, hoping to use a simple computer model to generate new musical sequences in his distinctive style.4 Around the same time, David Huron would start developing the first bespoke computational library for machine learning, Humdrum.5 In the 1990s, the psychologist Jamshed Bharucha published a series of influential articles arguing that computers can learn the music-theoretical concepts of scale degrees, chords, and keys simply by observing datasets.6The problem, however, was that twentieth-century machine-learning programs were downright lousy for making music. Motown would never use Raymond Scott’s machine. David Cope’s programs functioned more as compositional inspiration rather than the generators of final products. While Jamshed Bharucha confidently predicted that actual implementations of his theorized systems would be right around the corner, his architectures remained perpetually theoretical, never graduating from concept to reality.Early on, then, if you wanted to make music with a computer, you used explicit instruction. In contrast to machine-learned models that absorb knowledge from a dataset, these early musical AIs relied on hard-coded musical rules that governed their behaviors. In these explicit, rule-based systems, the composer uses their own knowledge to write code that creates music along specific guidelines—dictums like “write triads on downbeats”; “resolve your sevenths down”; “in this musical situation, choose from these notes.” A computer can then create music by referencing these reams of preprogrammed rules. Lejaren Hiller and Leonard Isaacson’s 1957 composition Illiac Suite, one of the earliest instances of AI-created music, would rely on this type of explicit logic. Work by subsequent computationally oriented artists, from H. J. Maxwell to Richard Boulanger to Iannis Xenakis, would incorporate versions of this concept.7It turns out that musical machine learning needed much more data and computing power to be successful. With the proliferation of online repositories and uploaded content, the 2000s and 2010s saw research based on the Million Song Dataset, the McGill Billboard Corpus, and the Corpus, that online into to academic began on machine learning in music, and like and would in the of computational this In computing power would and A that a to on my current would a in the at two would at a century for a in the like Leonard Meyer to the same of more data and computers and composers to more its in of a neural that relied on a dataset of to create new In the early the researchers and used a combination of AI, and human to on for his was a of the vast of were used to machine-learning the and the in they trained on of of music online and used huge of computing As the quarter of the twenty-first with an can these to generate musical with the of a to the and uses of these generative musical AIs take a the computational at how machine-learning and at the that these use the term Artificial Intelligence in a contemporary we’re a machine-learning and we’re technical as data, learning, neural or In this I’ll through this to a outline of how this technology an from David on music The on the tendencies of scale to to one and a greater for one scale to to and these to how that in a dataset of that a particular scale more often in the dataset, and the those be more and more In his Huron uses this to what of a might those they might is a simple of a computer system learning machine learning. Huron didn’t simply the that the and scale often one this was by all scale in the data, or the data used by a machine to its learning. and to how often scale to other The in then, model these behaviors. in this a is a of and a computer can use to or some The of the which scale sequences are which are model these The of could be used to generate new thereby the system as a generative would the more and the out scale along the scale sequences would be by a simple machine-learned generative this model is to and melodic not to be a generative would create that and to our and it any larger aspects of musical like and so the first of the of a simple that some of the of be for much of the scale For the first two scale and from scale to of these sequences are by the model has making of of scale and dataset of simply of these to as norms in these are as of the some to basic scale is but their on which is being used in this and My some basic ways these and their to in and could scale model coming from some machine-learning process. A model and scale in data of popular could result in this basic this model is of many machine-learned models of machine learning simply tallied scale in its data, of including the the of a type of model not simply rely on how often scale to one but rather on some of those scale to a these larger by to their computational one to it a of scale however, this In scale on whether they are a or it to these are programs of has with to this but neural networks the most in Artificial networks are so they use of to and of human this was explicit of the earliest in computing saw the and arguing in that computational networks the human could provide a for future While this research would over the coming most researchers and engineers would their emphasis from explicit with the human in of a and the twenty-first the neural has into an right out a simple of how these neural networks their Artificial networks of an along with and these capture how the is networks the new data, but the the machine-learning In I how a neural might the of scale to these various The a of the machine-learning as it might data from of In this the uses a series of scale from the first as its in to learn how that series of scale this point in its learning my has scale and can see these in the and the and the of has that scale in and it has these After many the model has the of what a human would a and a other around these In the those scale are all with a that that that same has to group scale using and as The behaviors of are in the those and the scale in the behaviors change with are the and to the scale of the the tendencies and with In for scale and the and the model the to be these the is how the model learns these behaviors. networks of a dataset as their and then to the as their In the in the scale of the melodic in this and the predicted scale is the the model that is it’s that the that this and predictions are drawn the and In the model the first as to the a change of a to the on scale that the final scale is that might it a With these in the model that the is most to be scale or As the this is and all the that to that would be this used for the neural The machine has concept of these how scale in has and scale one in and in and it these its computational is important to the of this of any music-theoretical knowledge that a might the model is a of what scale it this to the on my While has a of and and two with huge of and and of to model more the from that the by in for its used of in of The model was then trained on on the of in computational and to a neural for more musical predictions over were a with a vast series of and for one of these would learn to a series of scale the model some could predictions to might make and the predicted scale and the model, the more for more musical and greater engineers various to the of into the basic for in using its to that or data, like or in an or in music. use to of to of their dataset, like the of a or the of a as those used in use what are that and and the machine group its data into in or in music. then, are neural networks with that and their and to create more generative content. of neural networks are in their own to capture larger and more aspects of some dataset, and produce new using those a going on in the AI and in musical AIs in of and that I provide be outdated even by the this essay is are a few that like to in which generative use in music to be make a generative model using machine learning, you need of data, of a with in data a machine-learning model to be to trends and tendencies that it can use to generate new the data is with or the machine will learn to generate these of the actual of the Then, if any actual or tendencies in a dataset a machine-learning model be to it can use to generate new are a of particular that engineers musical that these In this I’ll outline of these to musical While at its on the five of a and those to on a are so many moving that can with those and or not and or not be or on the and on a or not be and all in how the music is could with the in a that are in with musical is of shows a from an for and of The right shows the of that musical from a The some from the most of the and are it some a and a many of the are in the and the the are as While the of the system is just as the a And, perhaps of some on the the right is with one of in the of of these how it is for a computer to and from music in other simply always to computer from musical is musical can be just as as to the of The of the that and the The then the coming from a and these using a a that shows on the and with greater As the and the the shows and at that The are all in the with you or a not but the that fundamental with more and and bright a of into the our and are at fundamental from their computers often this task I to the in this and it the in right the not as an but as or and that it two of that and as the then most of the on the to be of the The in other has and to a from the one as was the with computational these how it can be to musical from are shows an the of some and repositories with the use of a My shows a human would need to start to the of one of these online repositories in to its with in the and to the of need to start around the of the and to of would the were being on or on need to to the The of human of people write and all to the of these has some The of musical are on with those of and has the of the of the is to be more in of than or the of the Project the would to the its the to the series are musical so I compared with other music and and that the same or more often than or in other that the of musical data in the the of is the knowledge and technology needed to create music. on and how you around of the to new music and it for people make music, be of it While in the can their or a on a of a of that has the to create new musical a of the for much of across many for has some hard and rules. going to make much and and of on of rather than in from its a a system like a neural will be to these and some music like it be an for systems that on and to be or to be or scale to to to be than surrounding I could problem is that musical norms for of The of of for the type of twentieth-century The of related over a musical by a and this so many can In the and ideas and the use and and both the the of and these were The could other or could been in perhaps with or the upward of the could been a could been A or of the of the I could on, but the point is are ways to and of it hard for an mind to the norms and musical of these the particular in musical of with it’s hard to from and and musical data is and you make a dataset by from or your AI might end the of these you the musical dataset into the of you your AI learns the of musical you might be your is for a computer to the you this essay, will be music that from music by these do they like rather than to over In and for the in the contemporary musical AIs often generate or but to into a and to data and of they to incorporate musical and into their neural some of these I one of the most into some music. shows some While music they do a of knowledge compositional I to “write some code that and that composition as a where the was in with the in The code particular and which were then into a a I could and in music a few that the AI has in with For it seem to and in music, as of do in the same as I with it often the with larger and of at one compositional writing particularly the The first of in a combination of and is simply not The has some of the musical to that of be in the same it has how to at how to the in a a from a that generated in the of The has a with a over a even a and his the counterpoint the simple the and the the to that this for modes with and and it to that both not be in the same While an the simply to A and even the and the counterpoint is of these the AI has from its how to start and end a on a how to and even the use of and in it to more musical to how to create or melodic how to a how to use in a musical AIs take a a of the ways their and out where they at the of even the most advanced systems over versions of the by these While at it to more and they is one by which a model can these however, and to simply and of musical from the data into some using as of on of often this the process of simply specific the AI is trained on use this from this a an essay, simply from its dataset with the bit of a of first a by the generative AI by the of is an the is for its of musical knowledge by the of a in its how a specific create a that from of has My is drawn from the the of in of for at the most the and AI researchers rely on to make their systems create music just how human these for is to data, by the most aspects of a musical into digital or by musical data not as but rather as a series of In the first an AI will the of a musical into the of code to capture that a that the music to its most basic an is of musical as it has musical as much as to make its In the AIs often a musical like a rather than actual the to like the I in and learns how those to never at its seem for AI not of any state-of-the-art AIs that on to human and all its We music as a of as of and then, the what for machine learning and what for our human and of these might in the long music will start if it AI can out much and than a human With a few a or a can provide you with a of music. In this is so that music is vulnerable to AI replacement, more so than other I to one important and music that specific content, that specific in the and even that particular or musical are not to point to ideas or concepts in and and it rules. The a of a and the of a all a specific type of you a to you a and it the has and the program is not its you an AI to a and that has a and it’s The has particular rules. I this a bit I that musical as as or is used as an or if your has a the that generated is and being as it the it’s hard for music to be A can be can be and a can be but a or right to how a be and it’s not how a music at in the same as you would the of a with is and even if a the result is of the in are is a that it will be to use generative AI in music than in other to write a and it a that to few this the of the and is a to that the human if a some music with some or it’s as music. rules and it’s hard for music to be music has a to be is I will see a of music used for and the if music to like and that will be a of AI of these are do not the to a social or and from a musical some of the basic that create and value music, that will be the for a computer ever to a few ways that value to human For music can and social and in the for is in the historical and by music social and the of a on the of a or in a can to this perhaps most music a on a or our in a of we’re using music to and process our as from to to music by or human is human out of an that has been over and over in will and in music they can or with its in music they with a and they it was by a machine rather than by will the of in which generative AI will human While might a uses for a it’s hard to AI as a for the human of the or being a of on a hard AI will musical simply it is not this is in many from to to is it with AIs are particularly While or or other can by and music this a of music by us as for a to end this essay with a bit of a musical AI, aspects of the AI in began this essay with the AI is the worst the technology will ever be,” and on of the The few years some innovations in generative AI music and AI will change the ways teach and will the of do and the to our this essay has that music has a and with to the of the music particular for generative AI, and generative AI particular for music. to to to the music that much of our AI has a long and a road today’s musical AI the worst it will ever In all I I’m excited to see the in this
Fetched live from OpenAlex and de-inverted. Abstracts are not stored in this database: the inverted indexes are 8.6 GB of the frame’s 9.3 GB of text, and the host has 13 GB free.
How this classification was reachedexpand
Full frame distilled prediction
Teacher imitationNot calibrated prevalence, not ground truth. Human validation pending. Learned from the 10,348 direct Codex labels and 10,348 direct Gemma labels. Candidate is the union of thresholded teacher heads; consensus is their intersection. These outputs are machine_predicted_unvalidated and are not human labels or direct frontier model labels.
Codex and Gemma teacher scores by category
| Category | Codex | Gemma |
|---|---|---|
| Metaresearch | 0.001 | 0.000 |
| Meta-epidemiology (narrow) | 0.000 | 0.000 |
| Meta-epidemiology (broad) | 0.001 | 0.001 |
| Bibliometrics | 0.000 | 0.002 |
| Science and technology studies | 0.001 | 0.001 |
| Scholarly communication | 0.000 | 0.000 |
| Open science | 0.003 | 0.001 |
| Research integrity | 0.000 | 0.001 |
| Insufficient payload (model declined to judge) | 0.000 | 0.000 |
Machine scores (provisional)
The two teacher heads of the student model, read on this work. A score orders the frame for review; it never asserts a category, and the validation status ships verbatim with every row.
Baseline scores from an immature model (maturity gate not passed, 7 training rounds). Scores rank; they never assert a category.
score_only:v0-immature-baseline · verbatim from the scoring run: score_only means the number may rank works, and no category label ships from itClassification
machine, unvalidatedMachine predicted; a candidate call from one teacher head, not a consensus.
How this classification was reached, model by model and score by score, is at the end of the page under "How this classification was reached".